The guide moves from the market context through operating models, architecture, security and law to cost-effectiveness, decision aids and three practical examples from the SME sector.
Local AI is not a product decision but an operating decision. An installed language model is not yet an enterprise AI. Only a knowledge base, permissions, security, quality and durably owned operations turn it into reliable business value.
A growing number of small and mid-sized companies are examining whether to run generative AI on their own infrastructure. The reasons are understandable: sensitive company data should not leave the organization, dependence on individual cloud providers should fall, AI should remain available without an internet connection, usage costs should become more predictable, and company-specific knowledge should be integrated in a controlled way.
A local installation does not solve these requirements automatically. Installing a language model on a server does not yet create a usable enterprise AI. What is also needed is a robust knowledge base, roles and permissions, interfaces, logging, quality testing, security measures, updates and responsible operations.
Running AI in-house shifts responsibility from the provider to your own organization.
What matters is not model size but performance in the specific business process.
It should be supplied through a controlled knowledge layer such as RAG.
Legal basis, purpose limitation, data minimization and data subject rights remain relevant.
Models, knowledge, dependencies and security have to be operated permanently.
Sensitive tasks run locally; compute-intensive or fluctuating tasks run externally, under control.
“Installing a model is comparatively easy. Building a reliable enterprise system is the real task.”
Artificial intelligence has arrived in the SME sector, but it is not yet embedded across the board in day-to-day operations. The Federal Statistical Office and the KfW paint a consistent picture, even though their figures are not directly comparable because of differing definitions, periods and survey methods.
Among the companies that had examined but not yet deployed AI, the barriers cited in 2025 were:
| Barrier | Share |
|---|---|
| Lack of knowledge | 72% |
| Uncertainty about the legal implications | 62% |
| Data protection and privacy concerns | 60% |
| Incompatibility with existing systems | 45% |
| Insufficient data availability or data quality | 44% |
| Costs too high | 32% |
The core problem rarely lies in the AI model alone, but in the interplay of technology, company data, accountability and the ability to operate. The relevant question is therefore not “Which language model should we install?” but: “For which process do we need which AI capability, which data may be processed, and who takes durable responsibility?”
The term “local AI” is often used loosely. For a sound decision, at least five operating models should be distinguished. They differ in control, effort, scalability and suitability for sensitive data.
The model runs on a notebook or workstation. Suitable for personal research in approved documents, confidential draft texts, and development and testing. Limits: hard to administer centrally, no shared knowledge base, limited performance, and the risk of uncontrolled point solutions.
The model runs on servers in your own building or data center. Suitable for sensitive internal data, stable user groups, high and predictable usage, offline or production environments, and companies with their own IT operations.
The system runs in an isolated cloud environment or on dedicated servers in a data center. Suitable for companies without their own server room, multiple sites, controlled scaling, external operational support, and requirements for EU hosting or defined data locations.
The model runs directly on a machine, in a vehicle, at a branch or on a mobile device. Suitable for speech recognition in the field, image analysis at equipment, local assistance without a stable connection, and time-critical evaluation.
Local models and external AI services are combined through shared control. Internal documents and personal data are processed locally, general market information via external models. A routing layer decides based on the task, the data class and the required model performance.
For many SMEs the hybrid architecture is the most practical target architecture: it combines data sovereignty for sensitive processes with performance headroom for demanding tasks.
Four terms are frequently confused. Distinguishing them is the prerequisite for realistically assessing effort, value and responsibility.
Generates text, analyzes input, classifies content. But it does not automatically possess current company knowledge, access rights to ERP or DMS, awareness of the approved version of a document, traceability, or professional responsibility.
Provides the interface and the workflow, for example as an internal knowledge assistant, service assistant, quotation support, site documentation, maintenance assistance or tender analysis.
Connects the model to approved data sources: document management, ERP, CRM, ticketing, project files, operating instructions, bills of quantities, maintenance records, rulebooks and product catalogs.
Keeps the system usable over time: updates, permissions, logging, quality testing, backups, monitoring, incident handling, version changes, security reviews and support.
“The model produces answers. Only the application, the knowledge and the operations produce business value.”
Seven reasons justify local or largely local operation. In each case they should be demonstrable, not merely plausible in principle.
Technical drawings, costings, customer files, personnel and contract documents, price lists. Prerequisite: the local system itself is adequately secured.
Site containers, production halls, workshops, vehicles, ships and port facilities, remote locations, temporary venues.
With suitable infrastructure, lower and more predictable response times, for example for voice assistance, machine operation, image inspection or real-time classification.
With high, stable load, in-house operation can become cost-effective, especially with many short requests or large internal document collections.
Production, security or administrative networks without internet access can host a local model within the environment.
An approved version can be kept stable for longer. In return comes the duty to organize security updates and later migrations yourself.
Open-weight models and interchangeable runtimes reduce lock-in to a single provider. Full independence still rarely results: dependencies remain on hardware, licenses, operating systems, drivers, open-source components, partners and support.
Five widespread assumptions lead to poor decisions in practice. They should be clarified before any investment.
| Assumption | Why it is not quite true |
|---|---|
| “Our data does not leave the company.” | Only if telemetry, updates, logs, interfaces, front ends and add-on services are configured accordingly. Otherwise error reports, update services, external embedding or search services, or monitoring will still transmit data. |
| “Then we are GDPR-compliant.” | The storage location alone does not decide this. Purpose, legal basis, permitted data categories, data minimization, deletion concept, access and rectification processes, roles, logging and, where applicable, a data protection impact assessment remain to be checked. |
| “A local model does not hallucinate.” | Incorrect or invented answers are a property of generative models, not a consequence of cloud operation. Local systems also need source citations, thresholds, professional sign-off, refusal to answer on a thin data basis, and human control points. |
| “After installation there are hardly any costs.” | Beyond hardware there are administration, power and cooling, updates, security reviews, backups, model testing, interface maintenance, support, contingency provisions and technical renewal. |
| “Open source means free and unrestricted.” | Often only the model weights are available. License terms can restrict use, redistribution, modification or certain scenarios. Every model therefore needs a documented license review. |
The German Data Protection Conference recommends defining fields of use, purposes, legal bases, responsibilities, technical safeguards and the review of results before deployment.
Local AI delivers its value where your own data and process knowledge are decisive. Every professional answer should be returned with its source, document version and approval status.
An internal assistant searches work instructions, assembly manuals, product data sheets, maintenance documents, risk assessments, quality specifications and approved templates.
The technician describes a fault by voice. The system structures equipment, fault pattern, actions taken, spare-part needs and next step, and suggests relevant documents or earlier cases.
Minutes from voice notes, daily site reports, deviations from project files, open points from meetings, bills of quantities. Without replacing professional, technical or legal sign-off.
Connects traffic management plans, traffic orders, conditions, setup and inspection records, material lists and internal check notes. The decision on securing a specific site stays with qualified staff.
Structure shift handovers, summarize deviation reports, classify inspection reports, locate work instructions, match fault patterns against approved cases.
Classify enquiries, extract requirements from emails, find references, propose quotation building blocks, prepare CRM entries. Prices and commitments remain subject to approval.
For some tasks an external service or a hybrid solution is more cost-effective or simply better suited. These cases should be deliberately excluded.
If very high compute is only needed occasionally, external infrastructure is usually cheaper, for example for one-off analysis of very large document collections, extensive image or video generation, or infrequent bulk processing.
Smaller local models can lag behind leading external models on demanding strategy, research or problem-solving tasks.
A local model does not automatically know current laws, new product information, market prices, news, security advisories or amended standards. Such content requires controlled data sources or external research functions.
Particularly critical are decisions affecting recruitment and personnel assessment, creditworthiness, insurance benefits, medical treatment, access to essential services or security clearances. Additional legal, technical and organizational requirements apply here.
Quotations, complaints, service commitments or legally relevant statements should not be sent without defined approval rules.
The EU rules for certain high-risk AI systems were staggered by the Digital Omnibus adopted in 2026: for standalone high-risk applications (Annex III) they apply from 2 December 2027, and for AI in regulated products (Annex I) from 2 August 2028.
The capability of smaller models has risen sharply. The Stanford AI Index reports that the gap between open and closed models on selected benchmarks narrowed from eight to 1.7 percent within a single year. At the same time, the price of inference at the former GPT-3.5 level fell by more than a factor of 280 between November 2022 and October 2024. Hardware costs recently fell by around 30 percent and energy efficiency rose by around 40 percent per year.
The best AI is the smallest one that reliably fulfils the process. Do not oversize infrastructure for years, keep models interchangeable, do not decide on public benchmarks alone, reassess regularly, and test with smaller model classes first.
| Dimension | Key question |
|---|---|
| Professional quality | Are tasks solved correctly using your own documents? |
| Response time | Is the speed sufficient for the process? |
| Concurrency | How many employees use the system at the same time? |
| Resource requirements | What hardware is needed under normal and peak load? |
| Operability | Can the model, license and software be operated over time? |
A robust model test comprises at least 50 to 100 real tasks: typical questions, difficult edge cases, incomplete input, contradictory documents, unanswerable questions, confidential content, manipulated documents and factually wrong assumptions by the user. Approve only once defined minimum scores are met.
The memory required for the raw model weights can be estimated approximately:
On top of this come requirements for the runtime, intermediate results, context memory, several concurrent users, embedding models, security components, and the operating system and reserves.
| Model class | Typical use |
|---|---|
| Small models | Classification, extraction, short assistance, single workstation |
| Medium models | Knowledge assistant, document analysis, service support |
| Large models | More demanding reasoning, complex generation, higher infrastructure cost |
| Several specialized models | Routing by task, data class and required performance |
This classification does not replace a load test. Quantization, context length, output length and concurrency can change the requirements considerably.
For the workflow, other figures also matter: time to first output word, response time with 5, 10 or 20 users, behavior with long documents, stability after several hours, error rate on interface calls, load during indexing of new documents, and recovery after a failure.
Company knowledge usually belongs in a controlled knowledge layer, not in the model itself. With retrieval-augmented generation (RAG), the system first retrieves relevant content from an approved data basis and passes it to the model together with the question.
Policies, project documents, product information, maintenance documents, contracts, rulebooks and frequently updated content.
Advantages: knowledge updatable without retraining, sources displayable, permissions at document level, outdated content removable, answers restrictable to approved sources.
Recurring output formats, domain language, classification, specific dialogue patterns, structured extraction, standardized workflows.
Less suitable for frequently changing facts, current product data, document versions, permission control and content that must be deleted on demand.
“RAG provides knowledge. Fine-tuning changes behavior.”
In many enterprise solutions the combination makes sense:
a suitable base model,
a controlled knowledge base,
clear process rules,
and, where needed, a limited domain adaptation.
The German Data Protection Conference notes that RAG can influence the data protection assessment but still requires a case-by-case review. Legal bases, data subject rights, protection against data extraction and the security of the connected knowledge base remain relevant.
A robust enterprise architecture consists of eight coordinated building blocks.
Web, mobile, integration into line-of-business systems, voice or telephony interface.
Sign-in via existing accounts, roles and groups, document permissions, MFA.
Decides on model choice, permitted data, tool calls, logs and blocks.
Model management, resource control, parallelization, version changes, fallback.
Import, text recognition, chunking, metadata, vector and full-text search, versioning, deletion.
ERP, CRM, DMS, file server, email, calendar, ticketing, production, project management, telephony.
Errors, latencies, usage, model version, data source, quality and security events.
Approved use cases, data classes, control points, escalation, quality thresholds, accountability.
No block carries the load alone. Only their interaction makes the AI productive and controllable.
Generative AI brings additional attack surfaces. The German Federal Office for Information Security points, among other things, to risks from manipulated input, unintended disclosure of information, faulty results, and attacks on models and data. NIST recommends capturing, measuring and treating risks across the entire lifecycle, from selection and development through operation to decommissioning.
A document or input contains hidden instructions intended to override the desired behavior.
The model discloses information the user is not authorized to access.
A compromised document influences later answers.
Model files, containers, libraries or extensions contain vulnerabilities or malicious parts.
A system with write access to email, files, CRM or ERP performs unwanted actions.
Prompts, answers, credentials or personal data are logged without protection.
Read access can be broad; write access must remain tightly scoped and traceable.
Before productive use, the following should be documented: processing purpose, affected groups of people, data categories, legal basis, storage and deletion periods, recipients and interfaces, technical and organizational measures, options for rectification and deletion, and the need for a data protection impact assessment. The European Data Protection Board notes that models trained on personal data cannot be treated as anonymous by default; the possibility of extraction must be assessed case by case.
For performance assessment, recruitment, staff scheduling, time recording, behavioral or qualification analysis, the data protection officer, HR and, where applicable, the works council should be involved early.
The obligation to ensure an adequate level of AI literacy among the employees involved has applied since 2 February 2025. The Digital Omnibus, finally adopted in 2026, adjusted several deadlines:
| Date | What applies |
|---|---|
| since 02 Feb 2025 | Duty to ensure sufficient AI literacy among staff (Art. 4) |
| from 02 Aug 2026 | Official supervision and most transparency obligations (Art. 50), such as labelling of chatbots and deepfakes |
| from 02 Dec 2026 | Provider marking of AI-generated content (Art. 50(2)) for existing systems; new prohibition of abusive image content |
| from 02 Dec 2027 | Obligations for standalone high-risk systems (Annex III) |
| from 02 Aug 2028 | Obligations for high-risk AI in regulated products (Annex I) |
For local AI this means: users must understand limits and failure modes, administrators need additional technical knowledge, and business units must be able to assess results; training should be documented. This whitepaper does not constitute legal advice.
A productive AI operation distributes responsibility across clearly named roles. If one is missing, the typical weak points arise.
Strategic objectives, risk appetite, budget, organizational anchoring, approval of major areas of use.
Process value, functional requirements, test cases, acceptance criteria, control points, professional sign-off.
Architecture, availability, interfaces, updates, backups, technical documentation.
Threat analysis, safeguards, vulnerability management, monitoring, incident response.
Legal basis, purpose limitation, data minimization, data subject rights, deletion, impact assessment.
Approved sources, document versions, currency, archiving, functional ownership.
Test data sets, error patterns, approval thresholds, regular re-testing, quality reports.
Operation without a named system owner, a business unit that does not take part in testing, uncontrolled document imports, the administrator as the sole quality reviewer, unrestricted write access for AI agents, and productive operation without fallback and shutdown procedures.
Has the license changed? Is the required language supported? Does the answering behavior change? Do existing prompts and interfaces still work? Do resource requirements change? Are there new security risks? Are the functional minimum scores still met?
Model name and origin, license, checksum, quantization, runtime, system prompt, security rules, connected knowledge sources, test results, approval date, responsible persons.
A productive system needs a previous approved model version, a saved configuration and knowledge indexes, a documented restart, the ability to switch off individual functions, and a manual replacement process.
A sound cost-effectiveness analysis captures the total cost of operation over the lifecycle, not just the initial purchase.
Server or workstation, GPU, storage, network changes, installation, integration, building the knowledge base, security review, piloting.
Power and cooling, hosting, administration, monitoring, backups, updates, quality testing, support, replacement hardware, security management.
Downtime, missing functional ownership, upkeep of outdated documents, model migration, dependence on individuals, low user acceptance.
| Cost block | Local / year | External / year |
|---|---|---|
| Hardware depreciation | €7,000 | – |
| Energy and infrastructure | €3,000 | – |
| Administration and updates | €12,000 | €6,000 |
| Support and security | €5,000 | €4,000 |
| Usage fees | – | €30,000 |
| Total | €27,000 | €40,000 |
These example figures are not market prices. Lower usage, higher availability requirements or additional administration alone can reverse the result.
Local operation becomes more attractive when usage is high and stable, a suitable internal operating team is in place, sensitive data is processed, the required model class remains manageable, and there are no extreme load peaks.
Rate each criterion from 0 to 3 (0 = not relevant, 1 = low, 2 = relevant, 3 = business-critical) and multiply by the weight. Negative weights argue against in-house operation.
| Criterion | Weight | Rating |
|---|---|---|
| Particularly sensitive information | 3 | 0–3 |
| Operation without internet required | 3 | 0–3 |
| Constant high usage | 2 | 0–3 |
| Very short response time | 2 | 0–3 |
| Integration into isolated systems | 3 | 0–3 |
| Own IT operating capability | 3 | 0–3 |
| High demands on leading model performance | −2 | 0–3 |
| Highly variable load | −2 | 0–3 |
| Frequently changing external information | −2 | 0–3 |
| Limited operating budget | −2 | 0–3 |
Local or largely local operation should be examined in depth.
A hybrid architecture is often the best starting point.
A controlled external service or a private cloud is likely more cost-effective.
The matrix does not replace a data protection, security or cost-effectiveness review. It serves as a pre-selection.
From the use case to a sound operating decision, without starting from “We need a local language model”.
A concrete bottleneck, not a technology wish.
Often 10 to 20 users to start.
Approved documents, test data, anonymized cases.
Public, internal, confidential, specially protected.
At least 50 real tasks and edge cases.
Two model classes plus an external reference.
Owner, version, validity, permission, deletion rule.
Access, logging, backup, tamper protection.
Time saved, quality, error types, acceptance, cost.
Local, private cloud, hybrid or not yet.
Each fulfilled point counts as one. The check assesses organizational maturity, not technology alone.
Process described · value measurable · user group known · a business unit takes ownership.
Sources known · documents have owners · version and validity recognizable · permissions documented · outdated content removable · personal data identified.
Model performance tested · concurrency and response time measured · interfaces defined · backup and restore provided · monitoring planned · fallback procedure in place.
Threats assessed · write access limited · vetted model and software sources · logs without unnecessary sensitive content · incidents detectable and manageable.
Recorded in the AI register · data protection and information security involved · users trained · control points defined · quality criteria documented · an owner for updates · errors and complaints reportable.
Situation. A technical service provider with several field teams holds maintenance manuals, machine files, fault reports, spare-part information, email threads and the experience of individual technicians. The information is scattered across file servers, the DMS and various folder structures.
Target picture. A local service assistant should:
structure the reported fault pattern,
locate relevant documents,
show similar earlier faults,
suggest inspection steps,
prepare a service report.
Local language model, RAG knowledge base, sign-in via company accounts, access based on existing permissions, no automatic ERP changes, and source display on every professional answer.
The technician decides on diagnosis, spare parts and the work carried out. The AI provides suggestions and references, not binding determinations.
Time spent searching for documents, completeness of service reports, share of usable suggestions, functional errors, and acceptance in the field.
Situation. Site managers and project leads document site progress, obstructions, additional works, material use, acceptances, open points and photos. The information arises as voice notes, messenger messages, emails or handwritten notes.
Target picture. A local AI application should structure voice notes, assign entries to a project, flag missing mandatory information, prepare a daily site report, extract open points and deadlines, and locate relevant project documents.
| What the system does | What the system explicitly does not do |
|---|---|
| Structure and assign voice notes | no legal assessment of a variation |
| Flag mandatory entries and open points | no sending of service commitments |
| Prepare the daily site report | no changes to billing data without approval |
| Locate project documents | no replacement of review by site or project management |
Fewer media breaks, more complete documentation, faster project handovers, better findability of facts, and less dependence on individual knowledge.
Situation. A mid-sized company runs a DMS, an ERP, several file servers and many process descriptions. Employees frequently search for the current template, the valid work instruction, the responsible person, the approved product information or an earlier project decision.
Target picture. A Company Brain connects the approved sources with a local or hybrid language model. Every answer includes the source used, the document version, the approval status, the responsible unit and the time of the last update. Where the data basis is missing, the system does not answer speculatively but points to the responsible office.
Local AI can be a sensible building block for small and mid-sized companies. It offers advantages in data sovereignty, offline capability, integration and predictable usage. These advantages only materialize, however, once five prerequisites are met: a clearly defined business use case, a maintained and permission-controlled knowledge base, suitable infrastructure, defined security and governance procedures, and durably owned operations.
KrambergAI works with you to assess suitable use cases, data and protection needs, local, private and hybrid architecture variants, model and infrastructure requirements, operating and security requirements, the rough cost-effectiveness and the sensible scope of a pilot project. The result is a concise decision document with an architecture proposal, prerequisites, risks and a pilot plan.
Federal Statistical Office (Destatis) – Use of ICT in enterprises 2025.
26% of companies with 10 or more employees used AI in 2025 (small 23%, medium 36%, large 57%); the main barriers
were lack of knowledge, legal uncertainty and data protection.
destatis.de
KfW Research – Use of artificial intelligence in the SME sector (Focus on Economics No. 533, February 2026).
Special analysis of the KfW SME Panel: around 780,000 SMEs using AI (20%), construction 8%, R&D-active
companies 53%.
kfw.de
Eurostat – Use of Artificial Intelligence in Enterprises 2025. EU-wide data on adoption, areas of use and barriers.
ec.europa.eu/eurostat
Stanford Institute for Human-Centered AI (HAI) – AI Index Report 2025.
Gap between open and closed models on selected benchmarks narrowed from 8% to 1.7%; inference cost at GPT-3.5 level
fell by more than a factor of 280 between November 2022 and October 2024.
hai.stanford.edu
Federal Office for Information Security (BSI) – Generative AI models: opportunities and risks for industry and public authorities.
bsi.bund.de
German Data Protection Conference (DSK) – Guidance on the data-protection-compliant use of AI applications, including with reference to the RAG method.
datenschutzkonferenz-online.de
European Data Protection Board (EDPB)
– Opinion 28/2024 on data protection aspects of AI models. Models
trained on personal data cannot be treated as anonymous by default.
edpb.europa.eu
European Commission / Council of the EU
– EU AI Act and Digital Omnibus. On 29 June 2026 the Council gave its
green light to the staggered new application deadlines (Annex III from 2
December 2027, Annex I from 2 August 2028).
digital-strategy.ec.europa.eu · consilium.europa.eu
National Institute of Standards and Technology (NIST) – AI Risk Management Framework. A framework for governance, measurement and treatment of AI risks across the lifecycle.
nist.gov
The studies cited use different populations, company sizes, periods and methods. Percentages are therefore not directly comparable. Together they show growing and accelerating AI adoption alongside persistent gaps in knowledge, law, data quality and operations. This whitepaper is intended for professional orientation and does not constitute legal, data protection or information security advice.