Skip to content

Custom AI that ships with tests.

LLM features built into your product or operations stack with provider benchmarks, golden cases, traceability and an Exit Kit your engineers can own.

AI delivery console

Integration state ready for review

client-owned

Provider router

Anthropic Claude

Tool use, long-context reasoning

benchmarked

Boundary: not-for-training inputs

Azure OpenAI

EU tenant and Microsoft stack

ready

Boundary: EU-region option

Release gate

The prompt cannot ship alone.

Grounded answer

blocks release

>= target score

JSON contract

blocks release

0 parse failures

Refusal policy

owner sign-off

review sample

4

provider families

Claude, OpenAI, Azure OpenAI, Bedrock and open-weight options benchmarked before build.

CI

eval gate

Prompt and tool changes are judged before they reach the application.

trace

observable calls

Cost, latency, prompt version and outcome tags are visible per request.

repo

client-owned

Prompts, eval cases, runbooks and provider config live where your engineers can own them.

Provider benchmark

A side-by-side model decision record before a provider is wired into your app.

Golden cases

Real examples from your workflow, turned into repeatable checks for quality and safety.

Prompt repo

Prompt versions, tool schemas and output contracts committed as code.

Trace and alerting

Langfuse-style visibility for cost drift, latency drift, parse failures and provider errors.

The proof is in the workbench, not the pitch deck.

This is the shape of the artefacts that sit behind a production AI feature: route choices, release gates, trace fields and handover material. No fake customer data, no demo-only console.

Screenshot artefact

Provider switchboard

Anthropic Claude

Tool use, long-context reasoning

benchmarked

Azure OpenAI

EU tenant and Microsoft stack

ready

OpenAI API

Fast iteration and broad model coverage

optional

AWS Bedrock

AWS estates and vendor split

optional

Release proof

Eval gate report

Grounded answer

>= target score

blocks release

JSON contract

0 parse failures

blocks release

Refusal policy

review sample

owner sign-off

Latency and cost

p95 + budget cap

alerted

Release rule

The app can only ship when the use-case owner accepts the gate, the failure cases are named and the operator has a recovery path.

Operational proof

Trace and handover

ai.run.trace

prompt_version

ai.claims.v14

provider_route

azure_oai_eu -> claude

eval_tag

support_triage_golden

cost_guardrail

budget cap active

owner

product + engineering

Runbook

failure modes

Exit Kit

transfer-ready

Owner map

named

Budget

guarded

From use case to owned AI system.

Custom AI work fails when the model is treated as the product. The product is the contract around it: data, tools, evals, traces, security and handover. Select a state below to inspect what changes and what proof should exist.

Select a delivery state

The panel updates with the proof that should exist.

decision record

Name the decision boundary

The work starts with one workflow, one user action and a clear list of things the AI layer is not allowed to decide.

Use case

Provider route

Gate and trace

Exit Kit

Active artefact

ai-decision-record.md

Use-case boundary, failure modes, data residency notes and acceptance criteria.

Acceptance checks

Allowed action named
Human review rule set
Data path approved

Interaction model: each state maps to an artefact, a proof standard and a release check. That is the difference between a demo and a system someone can own.

The AI layer gets engineering discipline around it.

The service covers the parts most AI pages skip: provider selection, prompt and tool contracts, evals, telemetry, security notes and handover.

Boundary

AI product boundary

The feature is scoped around a real workflow, not a demo prompt. We name the user, the allowed actions, the failure modes and the evidence needed for launch.

Use-case decision record
Data boundary and residency notes
Go or no-go criteria
Routing

Provider router

The model vendor sits behind a route you control. That keeps provider changes, EU-region choices and cost decisions out of the application surface.

Claude, OpenAI, Azure and Bedrock routes
Fallback and retry policy
Cost-per-task model
Contracts

Prompt and tool system

Prompts, tools and output contracts are treated as product code with review, versioning and rollback instead of hidden strings in application logic.

Strict schema design
Tool input validation
Prompt-injection hardening
Release gate

Eval suite

Your real cases become a release gate. The goal is not to make the AI look clever, it is to catch regressions before users do.

Golden cases from real work
CI check for prompt changes
Human review where scoring is not enough
Operations

Production observability

The integration ships with the telemetry an operator needs: prompt version, route, tokens, cost, latency, error class and outcome tags.

Langfuse-style traces
Budget and latency alerts
Failure-mode runbook
Transfer

Security and handover

Data handling, access, audit logging and owner notes are documented so the AI layer can survive security review and vendor transfer.

GDPR and EU AI Act risk notes
Secrets and access boundary
Exit Kit handover pack

What your engineers receive at the end.

The engagement should leave a repo, a release gate and an operator path, not a mysterious prompt that only the vendor understands.

ai-decision-record.md

Artefact

Product + engineering

Owner

Use-case boundary, risk class, provider choice, acceptance criteria.

prompts/

Artefact

Engineering

Owner

Versioned prompts, tool schemas, output examples and rollback notes.

evals/

Artefact

Engineering + reviewers

Owner

Golden cases, CI command, scoring rules and review sampling plan.

observability/

Artefact

Operator

Owner

Trace fields, alert thresholds, dashboard spec and escalation route.

exit-kit/

Artefact

Client

Owner

Runbook, owner matrix, provider config and transfer brief.

Fixed stages, visible decisions.

The first decision is whether the use case deserves a custom build at all. If it does, the build moves through a narrow set of gates before production.

Phase 01

Scope + Model Selection

Scope gate

Weeks 1-2

Use-case scoping workshop, success-criteria definition, ROI modelling, model + provider benchmark on representative data. Output: architecture-decision record + phased plan.

Owner

AI Architect + Founder

Phase 02

Build + Eval Harness

Build gate

Weeks 3-8

Prompt + tool design, evaluation harness, provider abstraction, production plumbing (caching, retry, fallback, observability). Feature integrated into your application with CI gates.

Owner

AI Engineer + AI Architect

Phase 03

Operate or Hand-off

Operate gate

Week 10 onward

Operate-tier monthly performance review, quarterly model refresh, continuous prompt + cost optimization. Essential clients receive runbook library + 30-day support and own day-to-day operations.

Owner

AI Engineer + Account Manager

Start small or build the production layer.

Pricing is scoped after discovery because AI delivery depends on data access, system boundaries, review requirements and the number of integrations.

SCOPING

Essential

Discovery + ROI model + provider benchmark before committing to a full build engagement.

Commercial shape

Scope-quoted

Use-case scoping workshop
ROI model with break-even analysis
Provider benchmark on representative data
Architecture-decision record
Phased build plan with measurable milestones
Start with a spike
MOST POPULAR

Operate

Building a custom AI feature into your existing application with provider portability and CI-gated quality.

Commercial shape

Scope-quoted

Everything in Essential + full build engagement
Prompt + tool library committed to your repo
Evaluation harness with golden dataset
Production plumbing (caching, retry, fallback, observability)
Provider abstraction layer for portability
Monthly performance + cost review
Scope production build

Sovereign

Mission-critical AI features in regulated industries needing EU sovereignty, custom evals, and 24/7 incident response retainer.

Commercial shape

Scope-quoted

Everything in Operate
EU Data Boundary verification across all providers in stack
Self-hosted open-weights option as fallback
Custom evaluation framework for your specific compliance regime
24/7 incident response retainer with named on-call AI engineer
Quarterly red-team + safety review
Dedicated AI Architect + monthly executive review
Discuss operating model

Clear scope keeps the build honest.

The AI layer can be built quickly only when the surrounding product work is named. If the project needs a full SaaS build, training run or mobile app, that is a different scope.

Inside the scope

Use-case scoping + success criteria
Model + provider benchmarking
Prompt + tool design with version control
Evaluation harness + golden dataset
Production integration with caching / retry / fallback
Cost + latency observability

Outside the scope

Full SaaS product build (we build the AI layer; you own the product)
Custom model training from scratch (we use foundation models)
Mobile-app development (we integrate into existing apps)
24/7 model performance monitoring (separate retainer)

The route is chosen by the constraint.

Data residency, latency, cost, tool use and review requirements decide the provider route. The integration is built so that choice can change later.

Anthropic

Claude for tool use and reasoning checks.

Azure OpenAI

EU-region route for Microsoft-first tenants.

OpenAI

Broad model coverage behind the router.

AWS Bedrock

AWS route when the estate lives there.

Open-weight

Self-hosted option when data cannot leave.

Langfuse

Self-hosted trace and quality telemetry.

Microsoft tenant touchpoints

Microsoft AzureMicrosoft 365

Compliance mapping is evidence-oriented. It names guidance and controls used to shape the work, not a certification claim.

DORA evidence trace
ICT risk + resilience
NIS2 control trace
Art. 21 evidence
NIST CSF 2.0 trace
Public outcome map
Evidence-backed controls
Source + owner + test

Anthropic

API Partner

Direct Claude API access with Enterprise volume agreements and EU regional deployment.

Microsoft Azure OpenAI

Solution Partner (onboarding)

Azure OpenAI EU + Customer Lockbox for sovereignty-constrained tenants.

AWS Bedrock

Bedrock Partner Practice

Multi-model access (Claude, Llama, Mistral, Cohere) in EU regions.

Langfuse

Self-Hosted Observability

LLM trace + cost + evaluation backbone, self-hosted in your perimeter.

Serious AI questions need serious answers.

The section below is written for technical buyers, security owners and founders who do not want a demo that turns into operating debt. It names the limits as plainly as the delivery artefacts.

ScopeIs this a chatbot, an agent, or something else?

Not by default. The engagement starts with a workflow or product surface, then chooses the right interface. That may be a chat panel, a background classifier, a document workflow, a retrieval assistant, a tool-calling agent or a small internal console.

The decision is written into the AI decision record before build. If a chat UI is the wrong interface, it does not get built.

Evidence to expect

AI decision record
User action boundary
Acceptance criteria
ProviderWhich model provider do you prefer?

No provider gets picked by habit. Claude, OpenAI, Azure OpenAI, Bedrock and open-weight options are compared against your real cases, latency target, cost ceiling and data residency requirement.

The build uses a provider route, not hard-coded vendor logic. A provider change should be a controlled config and eval event, not a rebuild.

Evidence to expect

Provider benchmark
Cost-per-task model
Fallback route
DataDo you train a model on our data?

Usually no. Most production work uses hosted inference, retrieval, tool calls and evaluation. Custom model training is a separate scope and is out of scope for this service unless explicitly agreed.

If your data cannot leave your infrastructure, the route can be designed around self-hosted or tightly scoped inference. That decision affects cost, latency and maintenance, so it is made during scoping.

Evidence to expect

Data boundary note
Provider terms review
Residency route
ResidencyCan the AI route stay in the EU?

Often, yes, but it depends on model availability and the provider route at the time of scoping. Azure OpenAI EU regions, Bedrock EU options and self-hosted open-weight models are all valid routes when they fit the workload.

The page does not claim a blanket residency guarantee. The deliverable names the route, region, data class and remaining gaps.

Evidence to expect

Region note
Data class map
Residual risk log
QualityHow do you reduce hallucinations?

There is no honest zero-hallucination promise. The control is a bounded task, grounded data, strict outputs, retrieval checks, refusal behaviour and a golden-case eval set.

For safety-relevant tasks, human review stays in the path. The model can draft, classify, route or propose. The release rule states what it may decide alone, if anything.

Evidence to expect

Golden cases
Refusal checks
Human review rule
SecurityHow do you handle prompt injection and unsafe tool calls?

Prompt text is not treated as a security boundary. User-facing work gets threat modelling, tool allowlists, strict schemas, least-privilege credentials and server-side validation before any action is executed.

Retrieval sources are scoped. Tools only receive the fields they need. High-impact actions require review or an explicit approval step.

Evidence to expect

Tool schema
Permission boundary
Action approval rule
OwnershipWhat do we actually own at the end?

You own the prompt files, tool schemas, eval cases, route config, runbook and Exit Kit. The preferred shape is a client repository, with provider accounts and secrets in your tenant where possible.

The handover is designed so your engineer or next vendor can read the repo, replay the evals and continue delivery.

Evidence to expect

Client repo
Exit Kit
Runbook library
OperationsHow are cost and latency controlled?

The cost envelope is set before build. Production work includes per-call tracing, token and latency visibility, rate limits, budget guardrails, caching where it is safe and alerts for drift.

No open-ended agent loop goes live without a stop condition, a budget cap and an owner.

Evidence to expect

Cost envelope
Trace fields
Budget alert
InputsWhat do you need from us to start?

One target workflow, sample cases, access to the relevant APIs or exports, a technical owner, a product owner and a security contact. Non-production access is enough for the first pass.

The best starting point is ten to fifty real examples: good outputs, bad outputs and edge cases your team already recognises.

Evidence to expect

Sample case pack
API access list
Owner map
FitWhen would you recommend not building custom AI?

If a vendor feature already does the job, if the data is too weak, if no owner can review failures, or if the task needs legal accountability the business is not ready to hold.

A feasibility spike can end with a no-go. That is a valid outcome when it prevents a production system nobody can operate.

Evidence to expect

Go or no-go note
Build-vs-buy finding
Data readiness check
RiskDoes this cover GDPR and the EU AI Act?

It covers engineering evidence for the risk discussion: data flow, automated decisioning exposure, human review, logging, access boundaries and model route. It is not legal advice and it does not certify compliance.

Where GDPR Article 22 or EU AI Act obligations may be relevant, the risk is named early so counsel or the internal compliance owner can review it.

Evidence to expect

Risk classification
Review path
Control notes
IntegrationCan this integrate with our existing app and Microsoft stack?

Yes, when the system exposes a reliable API, event stream, database view or export path. Microsoft 365, Azure, Teams, SharePoint, line-of-business apps and internal services can all be valid integration points.

Fragile browser scraping is avoided for production. If it is used during discovery, it is labelled as temporary and replaced before launch.

Evidence to expect

Integration map
API permission list
Production route

Build the AI layer your team can operate.

Bring one workflow, one product surface or one internal operation. We will tell you whether it deserves a custom AI build, what the proof should be and what the first release would include.