Case studies
Results, not just deliverables
How we turn AI agents, automation, and custom AI into outcomes a CFO can verify — and how we document them.
The AI services industry has a case-study problem: anonymous logos, unverifiable percentages, and results that evaporate under a single follow-up question. We hold ourselves to a stricter standard. Every case study we publish names the workflow, states the baseline, and reports outcomes the client has signed off on — because if a result can't survive scrutiny, it isn't a result.
That standard also means this page grows slowly. We publish engagements only after systems have run in production long enough for the numbers to mean something, and only with the client's approval. In the meantime, the section below explains exactly how we measure — so you know what to expect from an engagement with us, and what to demand from anyone else you evaluate.
Case studies coming soon
We're documenting recent engagements. Want to discuss a similar project? Get in touch.
Our standard
How we measure results
Four rules that govern every number we ever publish — or claim in a sales call.
Baseline before build
Before any system ships, we record how the process performs today — hours spent, error rates, response times, revenue leakage — so 'better' has a number attached.
Instrumented from day one
Every system we deploy ships with its own measurement: tasks completed, accuracy against review, escalation rates, cost per task. Results come from dashboards, not recollection.
Honest attribution
We separate what the AI changed from what would have happened anyway — seasonality, staffing changes, process fixes — rather than claiming every uptick as ours.
Client-verified, or not published
A case study only appears here when the client has approved the numbers and the story. If we can't name the metric honestly, we don't publish the engagement.
The metrics that matter, by system type
Different AI systems earn their keep differently, so we anchor each engagement to the metric that fits. For AI agents and automation, that's hours of manual work retired, error rates versus the human baseline, and cost per completed task against the labor it replaces. For voice agents and chatbots, it's answered calls, booked appointments, resolution rate, and CSAT — not vanity "deflection" numbers that just mean customers gave up. For RAG and knowledge systems, it's answer accuracy on a golden question set and time saved per lookup. For machine learning, it's prediction accuracy against the method you use today, translated into dollars of decision improvement.
We agree on the metric in discovery, before a line of code is written. It becomes the definition of success in the statement of work — which is why our case studies, when they publish here, read like measurement reports instead of marketing copy.
Want to be a result we can publish?
Tell us what you're trying to achieve. We'll define the baseline, the metric, and the fastest path to moving it.