Contents
- Definition of open source AI and its key aspects
- Business benefits of implementing open source AI
- Risks and limitations of using open source AI
- Licences and legal aspects related to open source AI
- How to choose the right open source AI model and tools
- Deployment and maintenance of systems based on open source AI
- Use cases for open source AI in practice
- Is it worth using open source AI – decision recommendations
Share
Definition of open source AI and its key aspects
Open source AI, in practical terms, means access to the code, model weights and often also the training data, or at least detailed documentation. The most important question is whether the model can be run locally and modified, and the answer depends on the licence and on whether only the weights have been made available, or the inference code alone. Open models (e.g. Llama 3, Mistral, Qwen) allow you to deploy the solution on your own servers, which matters especially when the priority is control over whether data leaves the company. Closed models (e.g. GPT-4 in an API), by contrast, make it easier to get started quickly, but limit the possibilities in terms of changes, costs and the way data is processed.
“Openness” can exist on a spectrum, so before deployment you should explicitly check in the licence whether the model may be used in production and whether it can be redistributed. It is worth bearing in mind that code licences (e.g. Apache-2.0, MIT, BSD) do not always match model weight licences, which can introduce additional restrictions. In image generation, OpenRAIL-style licences are often encountered; these permit use, but impose obligations relating to prohibited uses. From a deployment perspective, the licence can be just as important as the quality of the model’s responses.
The open source AI ecosystem is built around repositories and tools that streamline testing and deployment. The largest hub remains Hugging Face (models, datasets, spaces), and GitHub as well as container registries (Docker Hub, GHCR) also play an important role. For local starts without specialist knowledge, Ollama, LM Studio or text-generation-webui are often used; they download models and provide a local API. In the image space, the standard is Stable Diffusion (e.g. SDXL) with tools such as Automatic1111 and ComfyUI, while in audio, Whisper (ASR) and TTS models (e.g. Coqui TTS) are popular.
- 01Access to componentsCode, model weights, data
- 02Open modelsLocal deployment, control
- 03Closed modelsQuick start, fewer options
- 04Graduated opennessCheck the licence before use
Openness gives control and the ability to modify, but requires licence verification; closed solutions mean speed and limitations.
Business benefits of implementing open source AI
Open source AI brings the greatest business benefits when privacy, control and predictable costs are key at higher usage volumes. OSS is most often chosen because data can remain on‑prem or in a private cloud, without sending prompts to an API provider, which streamlines work with sensitive content (e.g. contracts, emails, customer tickets) alongside your own logging and retention policies. With a high number of requests, token charges in an API can exceed the cost of maintaining your own infrastructure, especially for 7B–14B models with quantisation. An additional advantage is reduced vendor lock‑in risk, because you can migrate between models (e.g. Mistral → Qwen) while keeping a compatible API on your side.
- Privacy and data control thanks to local deployment (on‑prem) or in a private cloud.
- Better TCO efficiency under steady, high load (e.g. thousands of conversations per day) and 7B–14B models with quantisation.
- No vendor lock‑in: the ability to switch models without rewriting the entire product.
- Simpler integration with internal systems (ERP/CRM, SQL databases, document repositories) and RAG with connectors (e.g. LlamaIndex/LangChain).
- The ability to work offline and at the edge (e.g. GGUF/ONNX formats) when the Internet is limited or unavailable.
The biggest competitive advantage is usually provided by combining the model with your own unique data via RAG and tailoring behaviour through fine-tuning (e.g. LoRA). Fine-tuning allows you to adjust the tone, style and procedural knowledge to processes (e.g. IT helpdesk, complaints handling, product descriptions), while RAG enables work on up-to-date documents and the knowledge base. Open source also makes auditability and reproducibility easier, because you can version the model, prompts, data and pipeline, and then reproduce results when response quality changes. In many business use cases, a RAG + solid 8B/14B model delivers a better result than a very large model without access to domain data.
Risks and limitations of using open source AI
Open source AI involves quality, operational and security risks that must be consciously managed before use in production. Open models can still hallucinate, especially when they do not receive domain context or are given too much generation freedom, so you cannot “trust the answers 100%”. In practice, this is mitigated through RAG, source citation, rule-based validation and testing on control datasets. Another challenge remains assessing quality in Polish, because benchmarks are often in English and do not always translate 1:1 to real-world cases (inflection, proper nouns, law).
When a model has access to documents or tools (tool calling), you should assume a prompt injection risk and implement a sandbox, permission policies, content filtering and secret separation (e.g. Vault). Equally important remains supply chain security: malicious dependencies, swapped checkpoints or models with a backdoor, which is why it makes sense to verify sources (e.g. HF verified), scan container images, use artifact signatures (cosign) and isolate the runtime environment. OSS also lacks a default SLA and support guarantee, which can be problematic in 24/7 systems if the organisation does not have its own MLOps or a paid support contract. On top of that come less obvious maintenance costs: logs, monitoring, updates, regression tests, the indexing pipeline and access control.
Hardware limitations and runtime stability can be a hard “stop” for larger models and high availability. Large models (e.g. 70B) require many GPUs or aggressive quantisation, which affects quality and speed, and as a rough guide 7B in 4-bit often fits into 8–10 GB VRAM, 13B into 12–16 GB, while 70B usually requires ≥ 48–80 GB VRAM or distribution. It is also worth taking output drift into account: a change in quantisation, runtime version (CUDA, drivers) or retriever can noticeably shift the responses. This risk is reduced through a test suite, versioning (model registry) and canary deployments with rollback.
- 01No 100% trustRisk of hallucinations, requires context.
- 02Mitigation methodsRAG, validation, control tests.
- 03Deployment securitySandbox, permission policies, separation.
Open AI models require conscious management of quality, security and context before production deployment.
Licences and legal aspects related to open source AI
Licences and law in open source AI determine whether a model can be used commercially, deployed in SaaS and redistributed legally. Licences such as Apache-2.0, MIT and BSD are usually commercially “friendly” and allow you to modify and sell the solution, subject to attribution requirements. By contrast, GPL requires the release of derivative code upon distribution, while AGPL extends this obligation to making it available over the network (SaaS), which many companies consider a significant risk for server-side components. In addition, model weights licences can differ from code licences and may introduce restrictions on use or redistribution (e.g. Llama Community License), which sometimes means a “source-available” approach rather than fully open source.
Before a production deployment, it is worth separately verifying the weights licence, the server code licence and compliance with company policy (e.g. a ban on AGPL), and documenting the decision jointly with the legal department and the security team. In the image generation area, licences such as OpenRAIL are often encountered, which in principle allow use but also impose obligations related to prohibited uses (e.g. deepfake without consent), and regardless of the licensing terms, the right to image likeness must be taken into account. Significant IP risks linked to training data are also important: whether the data was lawfully obtained and whether the model does not reproduce protected fragments, which usually requires a risk assessment, usage rules (e.g. a ban on generating book excerpts) and anti-plagiarism filters. In the EU, requirements concerning documenting AI use, risk assessment and informing users in certain cases (AI Act) are also growing, which is why transparency in customer service and content generation can be an important element of compliance.
GDPR is easier to comply with with an on-prem hosted LLM, but the obligation to have a lawful basis, minimisation, retention, data subject rights and a DPIA in high-risk cases does not disappear. Prompt logs may contain personal data and trade secrets, so logging and retention policies, masking and practices such as selective logs and PII redaction are needed. With self-hosting, responsibility for security and availability passes to the organisation, so in higher-risk areas approval processes (human-in-the-loop), disclaimers and limiting high-risk uses (e.g. medicine, law) without validation are key. This approach also makes it easier to organise accountability when the question arises of who is responsible for incorrect advice generated by the model.
How to choose the right open source AI model and tools
Choosing an open source AI model and tools usually works best when the starting point is the requirements of a specific use case and testing against real questions. In practice, what matters are: the language (including quality in Polish), context length, effectiveness in tasks (QA, extraction, code), inference cost and licence. The safest decision is the one you confirm with your own benchmark, not with rankings and claims in model cards alone. Only after the initial selection does it make sense to fit the rest of the stack (runtime, RAG, vector database) to the intended load and environment.
- Select 2–3 models for the language and tasks (e.g. instruct variants for chat and customer support) and verify the weights licence.
- Choose embeddings for RAG and assess retriever relevance (often they determine the quality of the answer).
- Configure the runtime for the hardware: CPU/edge → llama.cpp, GPU and high traffic → vLLM or Hugging Face TGI.
- Plan quantisation (4-bit/8-bit) and measure the impact on task accuracy, not just the “feel” of the chat.
Text models are most often selected from families such as Llama 3, Mistral (including Mixtral, when quality with MoE matters) and Qwen2, which often performs well in multilingual settings. In use cases such as customer support in Polish, it makes sense to test instruct variants and check whether the model handles polite forms and proper nouns correctly. When code and analysis are the priority, it is worth considering specialised models such as StarCoder2, Qwen2.5-Coder or DeepSeek-Coder (depending on the licence). In programming, open source can be very effective, but the result largely depends on the language, repository, prompt quality and unit tests.
In knowledge-based systems for business, quality is often determined by the RAG components, i.e. the embeddings and retriever, rather than the generative model itself. In RAG, embeddings such as bge-m3, e5-large, GTE and multilingual sentence-transformers models are commonly used, and issues such as “the model returns the wrong passages” often come down to poor embeddings or an unsuccessful split of documents into chunks. Depending on scale, you can use FAISS locally or deploy a server-based vector database, e.g. Qdrant (open source) or Weaviate, when filters and frequent document updates are needed. Quantisation also matters for performance and cost: 4-bit (e.g. GPTQ/AWQ/GGUF Q4) reduces VRAM usage, and 8-bit is often a practical compromise worth confirming in tests.
“Fair” model comparison requires a repeatable evaluation process, not one-off attempts in chat. The most reliable approach is your own benchmark: 100–500 real questions and an assessment of correctness, citations, style and response time. To make results comparable, keep prompts fixed, use the same documents in RAG and measure metrics such as accuracy, groundedness, latency and cost per 1k queries. This approach also makes it possible to quickly identify whether a change in runtime, quantisation or retriever improves or worsens quality.
Deployment and maintenance of systems based on open source AI
Systems based on open source AI are most stable to deploy when the architecture (on‑prem or private cloud) is chosen at the outset to meet control and scaling requirements. On‑prem provides maximum control, but it involves investment in hardware and cooling, whereas a private cloud makes scaling and automation easier. At the start, a pilot in the cloud (GPU on demand) often works well, and only once quality and costs have stabilised does it make sense to move to permanent infrastructure. This approach reduces the risk of investing in hardware before the model’s fit to the data and workload has been confirmed.
Production does not end when the model is launched, because inference and working with data require a full set of services around it. The minimal production stack includes an inference server (vLLM/TGI/llama.cpp), an API gateway, authentication, logs, monitoring and a data indexing pipeline. In practice, observability (e.g. Prometheus/Grafana), alerting and cost control prove crucial, because without them it is difficult to maintain quality and predictable operation. This is particularly important when the system is meant to run 24/7 and serve many users in parallel.
Maintaining quality requires constant oversight not only of resources (CPU/GPU), but also of the model’s behaviour and outputs. It is worth tracking, among other things, the refusal rate, detected PII, RAG effectiveness, conversation length and user complaints, and catching quality drops through regression tests and alerts from metrics. To reduce answer drift after changes, a model registry (e.g. MLflow, Weights & Biases) and versioning of prompts and RAG configuration are used, preferably with immutable artefacts, i.e. unchangeable ones (tags, SHA), and a clearly described rollback procedure. Fine-tuning (e.g. LoRA) makes sense when a consistent style, response formats (e.g. JSON) or procedural knowledge is needed that cannot be provided by documents in RAG alone.
Security and continuity of operation can only really be ensured once protection of secrets, resistance testing against attacks and scaling plans have been taken into account. Secrets (keys, tokens) should be kept in HashiCorp Vault or Kubernetes Secrets, and access to tools should be restricted by RBAC, so that actions in systems (e.g. ERP) pass through a service layer with authorisation and auditing. It is worth implementing red-teaming tests (prompt injection, jailbreak, data exfiltration) and response policies and content filters, especially in knowledge-based use cases. Under heavier traffic, performance is usually improved by batching, parallelism, caching and choosing GPU instances, and issues such as “it is slow” more often stem from a lack of queuing and server settings (max tokens, concurrency) than from the model’s “power” itself.
Use cases for open source AI in practice
In practice, open source AI is most often used where data control and adapting the solution to an organisation’s processes and content are key. A typical deployment is a company assistant working on documents (policies, procedures, instructions) in a RAG architecture, so that answers are based on retrieved contexts. In this scenario, it is important that the user can see where the information comes from, so the system should return links or source excerpts. This approach works particularly well where up-to-dateness and compliance with documentation matter more than the model’s “general knowledge”.
Open source AI also works well in customer service and helpdesk automation, where the model can classify tickets, suggest replies and fill in fields in a ticketing system (e.g. Jira Service Management, Zendesk). Such solutions usually do not replace consultants 100%, but they can shorten response times and lighten the load on the first line, especially when they operate in a suggestion mode with human approval. In document work, data extraction to JSON also has great value (e.g. from invoices, contracts, CVs or reports), especially when the format is irregular and difficult to capture with regex-type rules. To maintain extraction quality, schema validation (e.g. JSON Schema) and tests on edge cases are needed, not just on “nice” examples.
Open source AI is also sometimes the basis of semantic search and knowledge analysis, where embeddings make it easier to find information more effectively in emails, notes, knowledge bases or code repositories, including in the case of paraphrases. In audio, local transcription (Whisper) is often used, followed by generating summaries and task lists without sending recordings to external services, while quality depends, among other things, on the level of noise, colloquial speech and VAD settings. In image generation, the Stable Diffusion/SDXL ecosystem makes it possible to quickly prepare moodboards and creative variants, and tools such as ControlNet help preserve composition, although this requires checking the model licence, rights to trademarks and brand safety policies. Agentic workflows and tool calling are also increasingly common, but their use should take into account limitations (e.g. read-only mode, operation limits, approval of critical actions and full auditing of tool calls).
Is it worth using open source AI – decision recommendations
Open source AI is worth using when the priorities are data control and predictable cost at high query volumes. It usually makes sense when sensitive content is being processed, greater control over model behaviour is needed, or the aim is to reduce the risk of vendor lock-in. At the same time, open source is not always cheaper, because with low traffic maintaining your own environment may cost more than the API. When the key priority is the fastest possible time-to-market and high quality without an ML team from day one, closed models exposed via an API may be the more sensible choice.
The decision is easiest to make after a short, structured pilot based on measurements rather than feelings alone. A practical “quick path” involves selecting 2–3 models and testing them on real questions, then building a simple RAG on Qdrant or FAISS, and finally adding monitoring and security policies. The absolute minimum for a deployment to make sense includes quality tests, access control, logs and a source citation mechanism in knowledge-based use cases. When assessing ROI, operational and qualitative metrics are helpful, such as reduced handling time (AHT), an increase in first-contact-resolution, lower cost per ticket, and the accuracy/groundedness of answers.
Open source AI also requires organisational decisions, because responsibility for security and availability largely shifts to the deploying side. A structured approach includes a governance model: a model owner, an update cycle, a change approval process and usage rules (allowed/forbidden), as well as the ability to quickly turn the feature off. At the same time, it is worth preparing a plan B for outages and degradation, for example switching to a smaller model, a no-generation mode (search only) or routing to a human. From the competence side, to reach stable production, you usually need roles covering backend (API), DevOps/Kubernetes, someone for data/RAG and an owner of content quality, although for a POC a smaller team is often enough than for 24/7 maintenance.
FAQ
Frequently asked questions
What are the biggest advantages of using open source AI in a company?
It usually comes down to greater privacy, control over data and more predictable costs at higher usage volumes. In addition, you can deploy the model locally, without sending prompts to the API provider.
Can open source AI be run locally on your own servers?
Yes, many open models can be deployed on-prem or in a private cloud. This matters especially when data should not leave the company.
Why is the licence so important when choosing open source AI?
Because it determines whether the model can be used commercially, modified and redistributed legally. You need to check the code licence and the model weights licence separately, because they may differ.
When can open source AI be cheaper than a closed API?
With a high number of queries, token fees in an API can exceed the cost of maintaining your own infrastructure. The article points out that this applies especially to 7B–14B models with quantisation.
What improves response quality the most in open source AI?
In many use cases, the key is combining the model with your own data via RAG and adapting its behaviour through fine-tuning. The article emphasises that often a better result comes from RAG + a solid 8B/14B model than from a very large model without domain data.
What risks should be taken into account when implementing open source AI?
Models can still hallucinate, and when they have access to documents or tools there is a risk of prompt injection. There are also supply chain security issues, no default SLA, and the costs of maintenance, monitoring and updates.




