Up front: why this roadmap starts with aborting
41 percent of German companies with 20 or more employees now use AI. A year earlier it was 17 percent, and a further 48 percent are planning or discussing adoption. The numbers come from Bitkom's March 2026 survey of German companies, and they mean: the question is no longer whether your company adopts AI. The question is whether adoption runs in a controlled way or as a gut-feeling project.
The same survey contains a second number that hardly anyone quotes: a third of AI users (33 percent) found that AI led to significantly higher costs. On the other side, 77 percent of users see their competitive position improved, and 66 percent want to expand their use.
In my experience, the difference between the disappointed and the satisfied is neither the tool nor the model: it is a roadmap that defines, before the start, what success is, what failure is, and when to abort.
Exactly there, the popular guides have a gap. I read the guides that rank at the top for this topic in the German market: Haufe, Mittelstand-Digital, Fraunhofer, Workday. All describe phase models, and some at least name typical mistakes. But none answers how you recognize that a pilot has failed. None lays its budget calculation open. And none connects the roadmap to the funding programs that exist for it.
This article closes all three gaps. You get: a checklist for when to postpone adoption, a roadmap in four phases, an open budget calculation with traceable line items, five measurable abort criteria, a funding check with real amounts, and the legal framework in three paragraphs.
Written from practice: at happycoding we build AI systems for mid-sized companies, and we have advised clients to abort. That is why this roadmap does not begin with the tooling but with the question of the conditions under which you should not start at all.
When you should postpone AI adoption
The most expensive mistake happens before the first prompt: adopting AI although the prerequisites are missing. Four situations in which I advise you to postpone.
No measurable process. If nobody can say how many cases come in per week and how long one takes, you cannot later judge whether the AI improved anything. Measure the current state first: case count, processing time per case, error rate. That costs one to two weeks of tally sheets, not software.
Without this baseline, every pilot is a blind flight, and every later success story is a claim. How you make processes ready for automation in the first place is covered in depth on our process automation page.
No reachable data. The most common sobering moment in first conversations: the knowledge the AI is supposed to work with lives in email inboxes, in Excel files on network drives, and in the heads of two long-serving colleagues. An AI can only use what is reachable by machine. If the data preparation is bigger than the actual project, then it is the actual project: consolidate first, automate second. That is not a rejection of AI, it is the right order.
No internal owner. An AI pilot without an internal champion dies after handover. You need one person with a real time budget who answers questions, spot-checks results, and represents the project internally. In our experience, four to eight hours per week during the pilot phase are enough. But those hours must be planned: 'someone will do it on the side' is the most common cause of death for working pilots.
A core-system migration is running in parallel. If your ERP switch is underway, the data foundation shifts away underneath the pilot. Then you lose twice: the pilot tests against a system that will soon no longer exist, and the migration gets competition for the same contact people. Migrate first, automate second.
Postponing does not mean cancelling: each of these four points can be fixed in weeks to a few months. But each one otherwise costs you money in the middle of the project, money an honest assessment at the start would have saved.
The roadmap: four phases from use-case selection to scaling
If the checklist above is green, the path looks like this. Four phases, each with its own timeframe and its own exit point.
Phase 1: choose the use case (one to two weeks). Take a process that meets three conditions: it occurs frequently, it follows recognizable rules, and an error can be corrected before it reaches the customer. Classics in mid-sized companies: classifying and pre-sorting incoming email, extracting invoice and delivery-note data, preparing quote drafts from existing data, answering internal knowledge questions. Choose a single process, not three: focus beats breadth, because otherwise you measure three half baselines instead of one whole one.
Phase 2: build the pilot (eight to twelve weeks). The range is an empirical value from our projects: it only goes faster if data and responsibilities really are settled up front. In this phase, a system emerges that works with real data from your daily operations, but still with a safety net: a human checks every output before it takes effect.
Technically, we usually build such pilots as workflows that run on your own infrastructure. What that looks like in practice is in our article on workflow automation with n8n on EU infrastructure.
Phase 3: evaluate (two weeks). Now the baseline pays off: you compare the pilot numbers against the current state, against criteria you wrote down before the start. What those criteria look like gets its own section further down, because they are the core of the whole roadmap. Important here: the evaluation is a meeting with a decision, not a quiet fade-out. It ends with one of three outcomes: scale, rework with a new timebox, abort.
Phase 4: scale (six to twelve months). Only now does deep integration pay off: connecting ERP and CRM, a permissions concept, monitoring, training the business departments, step-by-step expansion to neighboring processes. Six to twelve months sounds long, but it spreads across small expansion stages: each stage repeats the build-and-evaluate cycle in miniature, with the same abort criteria.
Remember: the four phases are not a waterfall but a chain of bets with growing stakes. Each phase buys you exactly the information you need to decide about the next one.
The budget framework: an open calculation
Hardly any guide names numbers, and those that do skip the derivation. So here is the calculation the way we set it up internally at happycoding. The person-days are empirical values from our projects; the day rates are typical market ranges for senior development in Germany: 1,000 to 1,400 euros per day. Plug in your own quoted rates, the structure stays the same.
The calculation, line item by line item
Scoping workshop: 2 to 3 person-days. Measure the process, check the data situation, cut the use case to size, write down the abort criteria. The result is a document you could use to commission any other vendor.
Pilot implementation: 10 to 15 person-days. Build the workflow, connect the model, test with real data, put in the safety net. The biggest block, and the only one that low-cost offers cover at all.
Integration: 3 to 5 person-days. Connecting email, ERP, or file storage, permissions, error handling. This block is underestimated most often, yet it decides adoption: a pilot that does not deliver its results to where your team works gets ignored.
Training and handover: 1 to 2 person-days. The team learns to check outputs, report errors, and know the limits. Without this line item, phase 3 measures not the AI but the confusion.
Added up: 16 to 25 person-days. Multiplied by the day rates, that makes 16,000 to 35,000 euros for a pilot including integration and training. On top come running costs, which are small in pilot operation: API costs by usage plus a small server, together typically a low three-digit euro amount per month (an empirical value from our pilot projects; the order of magnitude depends on volume).
Why 2,000-euro offers are a different product
Now for the contrast you will meet in your research: pilot offers of '2,000 to 10,000 euros' are circulating, for example in the cost guide by gerlinger.ai, justified by consulting experience, without an open calculation. I do not think the lower bound is a lie; I think it is a different product: at typical market rates, 2,000 euros is less than two person-days.
That is enough for a configured off-the-shelf tool with your documents. It is not enough for integration, data preparation, training, or data-protection documentation: exactly the blocks the same guide itself names elsewhere as cost drivers. These line items do not disappear; they surface later as a surprise.
For me, exactly that explains the Bitkom number from the beginning: for the 33 percent with significantly higher costs, it was often not the project that was too expensive but the initial offer that was smaller than the project.
An AI pilot costs as much as the person-days it needs. Every offer without a person-day calculation only shifts the costs to later.
Abort criteria: how you recognize that a pilot has failed
This is the section missing from every guide I read for this article. Five criteria with concrete thresholds that you write down before the start. The threshold values are suggestions from our project practice: adapt them to your project, but fix them before the pilot runs. Afterwards, nobody negotiates honestly with their own numbers anymore.
The five criteria with thresholds
1. Rework rate. How many AI outputs does a human have to substantially correct? Workday reports in its own report that around 40 percent of productivity gains are lost again to rework. If your team is substantially reworking more than about a third of the outputs at the end of the pilot phase, the checking eats the gain: abort or change the scope.
2. Error rate against the current process. The AI does not have to be error-free; it has to be better than or as good as your current process, at lower cost per case. That is why you need the baseline from the postpone checklist: if your team today handles three out of a hundred cases wrong and the AI eight, the case is clear, no matter how impressive the demo was.
3. Usage rate in the team. A pilot that fewer than half of the intended users use voluntarily after eight weeks has a problem that no model update will solve. Either the tool is built past the daily routine, or concerns in the team were never seriously addressed. Both are a reason to rebuild or abort, and you only find either if you actually measure usage.
4. Cost per case. Divide the running costs by the number of cases handled and compare with the cost of the manual process. If the AI case is above it after the pilot phase and no improvement through volume is in sight, scaling does not pay off. It is that sober.
5. Timebox exceeded. If no system is running with real data after twelve weeks, that is a result in itself: it usually means the data situation was worse than assumed. Do not extend quietly; go back to the postpone checklist and check which point was overlooked.
The hardest step: actually aborting
And then the hardest part: aborting when the criteria say so. The sunk-cost trap is especially treacherous in AI projects, because a new model has always just been released that supposedly changes everything.
So a change of perspective helps: a clean abort after the pilot costs you the pilot budget and delivers a measured process baseline, consolidated data, and the certainty of not sinking a six-figure scaling budget. That is not a failed project. That is a paid risk assessment, and it is the reason you will be faster on the next attempt.
Build, buy, or API: the technology decision
By phase 1 at the latest, the technology question is on the table: buy a finished tool, use a model via API, or build something of your own? The short answer for adoption: start as high up the shelf as possible. A pilot is supposed to answer a business question, not satisfy your curiosity about technology.
For the long answer we have a dedicated decision article that walks through the stages from off-the-shelf model to training your own: Training your own AI model: a decision guide for mid-sized companies. I deliberately will not duplicate the reasoning here, but three rules of thumb from it already help with pilot planning.
Company knowledge almost always means RAG. If the AI is supposed to answer questions from your documents, the standard architecture is a search across your own data plus a language model, not training your own. Why that is, and where the exceptions lie, is in RAG vs. fine-tuning.
Self-hosting is an arithmetic problem, not a matter of faith. Whether a self-operated model pays off depends on utilization, data-protection requirements, and operating effort. The full calculation is in What a self-hosted LLM really costs.
The architecture decision does not belong in the pilot. In the pilot you may be pragmatic: an API model with EU data processing, your data on your own infrastructure, human review in front of it. You answer the question 'do we build something of our own long-term?' in phase 4 with the real numbers from the pilot, not before with guesses. That saves you the most expensive wrong decision of all: an architecture for a need that, after measuring, does not exist.
Funding check: grants that make the pilot genuinely cheaper
The budget calculation above can be roughly cut in half in two German states. That is not an exaggeration; it is the respective funding rate. Both are German regional funding programs: they apply if your company operates in the respective state.
Mecklenburg-Vorpommern: a 50 percent grant, up to 50,000 euros per project. The digitalization funding via the Technologie-Beratungs-Institut (TBI) targets companies under 100 employees, but only in certain sectors: manufacturing, skilled trades, tourism. There is also a minimum project volume, staggered by company size, starting at 15,000 euros: that fits a serious pilot with integration well and a 2,000-euro experiment badly.
Brandenburg: BIG-Digital, a 50 percent funding rate, up to 250,000 euros for implementation. Plus up to 50,000 euros each for consulting and for training. The program thus covers exactly the line items in my budget calculation. Important for timing: the directive expires on June 30, 2027 (Germany's federal funding database, retrieved August 18, 2026), and the application must be submitted before the project starts.
Two honest caveats. First: funding applications are work in themselves; plan for a few weeks of lead time plus approval time, and schedule the project start behind it. Second: funding must never be the reason for the project. If the pilot does not pay off without the grant, the grant does not turn it into a good project, just a subsidized bad one.
As an accelerator for a project that passes the abort-criteria logic anyway, both programs are real money: a 30,000-euro pilot budget becomes a 15,000-euro own share.
Legal framework: what the AI Act means for your adoption
For most adoption projects at mid-sized companies: less drama than the headlines suggest. You still need to know three facts, because they determine what you document and what you label on your website.
The deadlines were postponed, not abolished. Since July 27, 2026, the omnibus regulation, Regulation (EU) 2026/1744, has been in force. It postpones the obligations for high-risk systems: Annex III (including AI in hiring decisions or credit scoring) applies from December 2, 2027, Annex I from August 2, 2028. Many older guides still cite the old dates. So for every text about the AI Act, check the publication date first.
The transparency obligations already apply. Article 50 has been applicable since August 2, 2026 and was left unchanged by the omnibus: when people interact with an AI or see AI-generated content, that must be recognizable. For an internal pilot with human review, that is usually quickly satisfied. For a customer-facing chatbot, it is a real requirement that belongs in the pilot person-days.
Typical pilots at mid-sized companies are not high-risk. Email sorting, document extraction, internal knowledge search: none of that falls under Annex III. But you have to document this classification, not just assume it, and with HR topics the line is reached quickly.
Deliberately, I will not give you more here: I am a developer, not a lawyer, and legal advice belongs in other hands. What you as a buyer should settle contractually and organizationally is in our article EU AI Act for software buyers: what applies from August 2026. There you will also find the line where a deployer becomes a provider.
Next steps
The roadmap in this article works without us, and that is exactly how it is meant. If you want to start on your own, the order is: measure one process (case count, processing time, error rate), go through the postpone checklist, write down five abort criteria, and only then collect offers. These three pieces of preparation alone sort the vendors for you: whoever reacts nervously to your abort criteria is not planning with your success but with your budget.
If you would rather take the first step with us: that is exactly what our scoping workshop is for, the two to three person-days from the budget calculation above. At the end you have the measured process, the tailored use case, your abort criteria, a budget calculation following the pattern of this article, and a funding assessment for your region. With the result you can build with us or with anyone else: the document is yours.
The simplest first step: book a no-obligation intro call. 30 minutes, we look together at one concrete process from your company, and you get an honest assessment: whether a pilot makes sense now, what it would cost by our calculation, or whether you are better off postponing. We will tell you the last one openly too; you have read above why.
