Custom AI that ships with tests.
LLM features built into your product or operations stack with provider benchmarks, golden cases, traceability and an Exit Kit your engineers can own.
AI delivery console
Integration state ready for review
Provider router
Anthropic Claude
Tool use, long-context reasoning
Boundary: not-for-training inputs
Azure OpenAI
EU tenant and Microsoft stack
Boundary: EU-region option
Release gate
The prompt cannot ship alone.
Grounded answer
blocks release
>= target score
JSON contract
blocks release
0 parse failures
Refusal policy
owner sign-off
review sample
4
provider families
Claude, OpenAI, Azure OpenAI, Bedrock and open-weight options benchmarked before build.
CI
eval gate
Prompt and tool changes are judged before they reach the application.
trace
observable calls
Cost, latency, prompt version and outcome tags are visible per request.
repo
client-owned
Prompts, eval cases, runbooks and provider config live where your engineers can own them.
Provider benchmark
A side-by-side model decision record before a provider is wired into your app.
Golden cases
Real examples from your workflow, turned into repeatable checks for quality and safety.
Prompt repo
Prompt versions, tool schemas and output contracts committed as code.
Trace and alerting
Langfuse-style visibility for cost drift, latency drift, parse failures and provider errors.
The proof is in the workbench, not the pitch deck.
This is the shape of the artefacts that sit behind a production AI feature: route choices, release gates, trace fields and handover material. No fake customer data, no demo-only console.
Screenshot artefact
Provider switchboard
Anthropic Claude
Tool use, long-context reasoning
Azure OpenAI
EU tenant and Microsoft stack
OpenAI API
Fast iteration and broad model coverage
AWS Bedrock
AWS estates and vendor split
Release proof
Eval gate report
Grounded answer
>= target score
blocks release
JSON contract
0 parse failures
blocks release
Refusal policy
review sample
owner sign-off
Latency and cost
p95 + budget cap
alerted
Release rule
The app can only ship when the use-case owner accepts the gate, the failure cases are named and the operator has a recovery path.
Operational proof
Trace and handover
ai.run.trace
prompt_version
ai.claims.v14
provider_route
azure_oai_eu -> claude
eval_tag
support_triage_golden
cost_guardrail
budget cap active
owner
product + engineering
Runbook
failure modes
Exit Kit
transfer-ready
Owner map
named
Budget
guarded
From use case to owned AI system.
Custom AI work fails when the model is treated as the product. The product is the contract around it: data, tools, evals, traces, security and handover. Select a state below to inspect what changes and what proof should exist.
Select a delivery state
The panel updates with the proof that should exist.
Name the decision boundary
The work starts with one workflow, one user action and a clear list of things the AI layer is not allowed to decide.
Use case
Provider route
Gate and trace
Exit Kit
Active artefact
ai-decision-record.md
Use-case boundary, failure modes, data residency notes and acceptance criteria.
Acceptance checks
Interaction model: each state maps to an artefact, a proof standard and a release check. That is the difference between a demo and a system someone can own.
The AI layer gets engineering discipline around it.
The service covers the parts most AI pages skip: provider selection, prompt and tool contracts, evals, telemetry, security notes and handover.
AI product boundary
The feature is scoped around a real workflow, not a demo prompt. We name the user, the allowed actions, the failure modes and the evidence needed for launch.
Provider router
The model vendor sits behind a route you control. That keeps provider changes, EU-region choices and cost decisions out of the application surface.
Prompt and tool system
Prompts, tools and output contracts are treated as product code with review, versioning and rollback instead of hidden strings in application logic.
Eval suite
Your real cases become a release gate. The goal is not to make the AI look clever, it is to catch regressions before users do.
Production observability
The integration ships with the telemetry an operator needs: prompt version, route, tokens, cost, latency, error class and outcome tags.
Security and handover
Data handling, access, audit logging and owner notes are documented so the AI layer can survive security review and vendor transfer.
What your engineers receive at the end.
The engagement should leave a repo, a release gate and an operator path, not a mysterious prompt that only the vendor understands.
ai-decision-record.md
Artefact
Product + engineering
Owner
Use-case boundary, risk class, provider choice, acceptance criteria.
prompts/
Artefact
Engineering
Owner
Versioned prompts, tool schemas, output examples and rollback notes.
evals/
Artefact
Engineering + reviewers
Owner
Golden cases, CI command, scoring rules and review sampling plan.
observability/
Artefact
Operator
Owner
Trace fields, alert thresholds, dashboard spec and escalation route.
exit-kit/
Artefact
Client
Owner
Runbook, owner matrix, provider config and transfer brief.
Fixed stages, visible decisions.
The first decision is whether the use case deserves a custom build at all. If it does, the build moves through a narrow set of gates before production.
Phase 01
Scope + Model Selection
Weeks 1-2
Use-case scoping workshop, success-criteria definition, ROI modelling, model + provider benchmark on representative data. Output: architecture-decision record + phased plan.
Owner
AI Architect + Founder
Phase 02
Build + Eval Harness
Weeks 3-8
Prompt + tool design, evaluation harness, provider abstraction, production plumbing (caching, retry, fallback, observability). Feature integrated into your application with CI gates.
Owner
AI Engineer + AI Architect
Phase 03
Operate or Hand-off
Week 10 onward
Operate-tier monthly performance review, quarterly model refresh, continuous prompt + cost optimization. Essential clients receive runbook library + 30-day support and own day-to-day operations.
Owner
AI Engineer + Account Manager
Start small or build the production layer.
Pricing is scoped after discovery because AI delivery depends on data access, system boundaries, review requirements and the number of integrations.
Essential
Discovery + ROI model + provider benchmark before committing to a full build engagement.
Commercial shape
Scope-quoted
Operate
Building a custom AI feature into your existing application with provider portability and CI-gated quality.
Commercial shape
Scope-quoted
Sovereign
Mission-critical AI features in regulated industries needing EU sovereignty, custom evals, and 24/7 incident response retainer.
Commercial shape
Scope-quoted
Clear scope keeps the build honest.
The AI layer can be built quickly only when the surrounding product work is named. If the project needs a full SaaS build, training run or mobile app, that is a different scope.
Inside the scope
Outside the scope
The route is chosen by the constraint.
Data residency, latency, cost, tool use and review requirements decide the provider route. The integration is built so that choice can change later.
Anthropic
Claude for tool use and reasoning checks.
Azure OpenAI
EU-region route for Microsoft-first tenants.
OpenAI
Broad model coverage behind the router.
AWS Bedrock
AWS route when the estate lives there.
Open-weight
Self-hosted option when data cannot leave.
Langfuse
Self-hosted trace and quality telemetry.
Microsoft tenant touchpoints
Compliance mapping is evidence-oriented. It names guidance and controls used to shape the work, not a certification claim.
Anthropic
API Partner
Direct Claude API access with Enterprise volume agreements and EU regional deployment.
Microsoft Azure OpenAI
Solution Partner (onboarding)
Azure OpenAI EU + Customer Lockbox for sovereignty-constrained tenants.
AWS Bedrock
Bedrock Partner Practice
Multi-model access (Claude, Llama, Mistral, Cohere) in EU regions.
Langfuse
Self-Hosted Observability
LLM trace + cost + evaluation backbone, self-hosted in your perimeter.
Serious AI questions need serious answers.
The section below is written for technical buyers, security owners and founders who do not want a demo that turns into operating debt. It names the limits as plainly as the delivery artefacts.
ScopeIs this a chatbot, an agent, or something else?
Not by default. The engagement starts with a workflow or product surface, then chooses the right interface. That may be a chat panel, a background classifier, a document workflow, a retrieval assistant, a tool-calling agent or a small internal console.
The decision is written into the AI decision record before build. If a chat UI is the wrong interface, it does not get built.
Evidence to expect
ProviderWhich model provider do you prefer?
No provider gets picked by habit. Claude, OpenAI, Azure OpenAI, Bedrock and open-weight options are compared against your real cases, latency target, cost ceiling and data residency requirement.
The build uses a provider route, not hard-coded vendor logic. A provider change should be a controlled config and eval event, not a rebuild.
Evidence to expect
DataDo you train a model on our data?
Usually no. Most production work uses hosted inference, retrieval, tool calls and evaluation. Custom model training is a separate scope and is out of scope for this service unless explicitly agreed.
If your data cannot leave your infrastructure, the route can be designed around self-hosted or tightly scoped inference. That decision affects cost, latency and maintenance, so it is made during scoping.
Evidence to expect
ResidencyCan the AI route stay in the EU?
Often, yes, but it depends on model availability and the provider route at the time of scoping. Azure OpenAI EU regions, Bedrock EU options and self-hosted open-weight models are all valid routes when they fit the workload.
The page does not claim a blanket residency guarantee. The deliverable names the route, region, data class and remaining gaps.
Evidence to expect
QualityHow do you reduce hallucinations?
There is no honest zero-hallucination promise. The control is a bounded task, grounded data, strict outputs, retrieval checks, refusal behaviour and a golden-case eval set.
For safety-relevant tasks, human review stays in the path. The model can draft, classify, route or propose. The release rule states what it may decide alone, if anything.
Evidence to expect
SecurityHow do you handle prompt injection and unsafe tool calls?
Prompt text is not treated as a security boundary. User-facing work gets threat modelling, tool allowlists, strict schemas, least-privilege credentials and server-side validation before any action is executed.
Retrieval sources are scoped. Tools only receive the fields they need. High-impact actions require review or an explicit approval step.
Evidence to expect
OwnershipWhat do we actually own at the end?
You own the prompt files, tool schemas, eval cases, route config, runbook and Exit Kit. The preferred shape is a client repository, with provider accounts and secrets in your tenant where possible.
The handover is designed so your engineer or next vendor can read the repo, replay the evals and continue delivery.
Evidence to expect
OperationsHow are cost and latency controlled?
The cost envelope is set before build. Production work includes per-call tracing, token and latency visibility, rate limits, budget guardrails, caching where it is safe and alerts for drift.
No open-ended agent loop goes live without a stop condition, a budget cap and an owner.
Evidence to expect
InputsWhat do you need from us to start?
One target workflow, sample cases, access to the relevant APIs or exports, a technical owner, a product owner and a security contact. Non-production access is enough for the first pass.
The best starting point is ten to fifty real examples: good outputs, bad outputs and edge cases your team already recognises.
Evidence to expect
FitWhen would you recommend not building custom AI?
If a vendor feature already does the job, if the data is too weak, if no owner can review failures, or if the task needs legal accountability the business is not ready to hold.
A feasibility spike can end with a no-go. That is a valid outcome when it prevents a production system nobody can operate.
Evidence to expect
RiskDoes this cover GDPR and the EU AI Act?
It covers engineering evidence for the risk discussion: data flow, automated decisioning exposure, human review, logging, access boundaries and model route. It is not legal advice and it does not certify compliance.
Where GDPR Article 22 or EU AI Act obligations may be relevant, the risk is named early so counsel or the internal compliance owner can review it.
Evidence to expect
IntegrationCan this integrate with our existing app and Microsoft stack?
Yes, when the system exposes a reliable API, event stream, database view or export path. Microsoft 365, Azure, Teams, SharePoint, line-of-business apps and internal services can all be valid integration points.
Fragile browser scraping is avoided for production. If it is used during discovery, it is labelled as temporary and replaced before launch.
Evidence to expect
Build the AI layer your team can operate.
Bring one workflow, one product surface or one internal operation. We will tell you whether it deserves a custom AI build, what the proof should be and what the first release would include.