Work

The systems we build.

Client work is confidential by default — we name a client only with their written permission, and we don't publish numbers we can't stand behind. What we can show you is the shape of the work: the problems we take on and how we approach them. References are available on request during a serious conversation.

Document intelligence

Turning unstructured documents into decisions

Bills, invoices, applications, field photos — high-volume documents that a person currently reads one at a time. We build extraction and verification pipelines with confidence scoring and human-review routing, so the easy 90% is automated and the ambiguous 10% reaches a person with context attached.

The part that makes these systems trustworthy isn't the model — it's the evaluation loop. Every correction a reviewer makes feeds back into a measured accuracy baseline, so quality is a number that moves, not an impression.

Retrieval & knowledge

Making an organization's knowledge answerable

Policy manuals, program rules, support histories — knowledge that exists but can't be asked a question. We build retrieval systems that answer with citations back to source, refuse when the source doesn't support an answer, and log every response for audit.

The design constraint we hold hardest: a wrong answer delivered confidently is worse than no answer. Grounding and refusal behavior get evaluated as rigorously as accuracy.

Evaluation & migration

Answering "which model?" with evidence

Teams often arrive running the model they started with, unsure whether a newer or cheaper one would serve them better. We build an evaluation set from their production traffic, benchmark the candidates on accuracy, cost, and latency, and deliver a decision with the data behind it.

Sometimes the verdict is "switch and save." Sometimes it's "stay put — the migration cost outweighs the gain." Because no vendor pays us, either answer is fine with us.

Agentic workflows

Automating multi-step work with guardrails

Workflows where the AI has to plan, use tools, and act — processing a request end to end rather than answering a question. We build these with hard boundaries: deterministic code owns money, permissions, and irreversible actions; the model owns understanding and language.

That division isn't a limitation, it's the architecture. It's what lets an agent be useful on Monday and auditable on Friday.

Have a problem in this shape?

Tell us about it. If it's a fit, we'll walk you through comparable work and connect you with references.

Start a conversation