Skip to content
Menu

How we operate

Six pillars. One disciplined way of working.

AI implementation is a craft, not a product. This is the operating framework we bring to every engagement, built from eight years of deploying systems inside real businesses rather than theorising about them from a conference room.

The six pillars

Every AI vendor claims to be different. Most mean they have a different product. We mean we have a different operating model. These six pillars describe how we actually work: the structure, the discipline, and the commitments that make our implementations stick. Each of them has something behind it we can show you, which is the only reason any of them are worth printing.

01

Forward-deployed engineering

Two named people

Hands-on implementation team, embedded in your business.

Every engagement is staffed by two people: a forward-deployed engineer who builds and maintains the systems, and an account manager who owns the relationship and keeps implementation aligned with business outcomes.

The term "forward-deployed engineer" was popularised by firms like Palantir to describe technical implementers who work inside a client's environment. Not from a remote agency desk. Not via weekly video check-ins. They sit alongside the people whose work they are affecting. We adopted the same structure because AI implementation fails most often in the last mile, the gap between "the tool works" and "the tool is being used correctly inside this specific business." That last mile requires someone on the ground.

Pairing the engineer with an account manager is deliberate. The engineer knows the tools, the architecture and the code. The account manager knows your goals, your team's adoption patterns and your reporting needs. Most AI shops send one person who does both badly. We send two who each do one well.

Embedding is also what makes the work attach to something. Every system we build is tied to a goal someone on your team already owns and is already measured on, not to a goal we invented in order to justify the build. If nobody inside the business is accountable for the number a system is supposed to move, that is a signal the system should not be built yet.

Practically: you get a named engineer and a named account manager. Both attend your first call. Both are reachable during business hours. Both attend monthly reviews. Neither disappears when a project phase ends.

What this means for you. When something breaks, you know who to call. When something needs adjustment, two people who already understand your business show up to handle it. No rotating junior staff. No handoffs to strangers.

02

Audit before action

Two week process

We inventory before we implement. Always.

Before we recommend a single tool, workflow or automation, we map what you already have. The existing stack. The active subscriptions. The automations quietly running, or quietly broken. The metrics nobody is looking at.

Most AI vendors show up selling. We show up asking questions. The audit is a two week process where we sit inside your business, review your tools, interview your team, and produce a written inventory of what exists, what it costs and what it produces. You see the output before anyone changes anything.

This matters because the problem is rarely a lack of tools. It is tool sprawl. We regularly find businesses paying for three chatbot platforms, two automation tools and a writing subscription nobody has opened in months. Sometimes the audit alone produces a return before we deploy anything new.

Audits produce three deliverables: a map of your current AI stack with cost attribution, a prioritised list of the three highest-return opportunities for new implementation, and a list of tools we would recommend cutting. You decide which to pursue.

What this means for you. You will not find yourself six months in wondering why you are paying for something that does not work. The audit creates a baseline we both reference every month. If a tool is not earning its place, we cut it, regardless of whether we recommended it.

03

Measurement-first implementation

Baseline before build

Every system ships with a KPI. No measurement, no deployment.

Every system we deploy has a defined success metric before it goes live. Hours saved per week. Leads captured per month. Response time reduced. Revenue attributed. If we cannot measure the outcome, we do not ship the implementation.

The reason is visible in the market data. More than a quarter of organisations using AI, 26 percent, do not know whether it changed their profitability over the past year. Only 36 percent say it improved.

Source: McKinsey & Company, The State of AI in 2025, as published in the Stanford HAI Artificial Intelligence Index Report 2026, April 2026. 1,993 respondents across 105 countries; fielded 25 June to 29 July 2025.

That is not a failure of the models. It is the absence of measurement infrastructure. A chatbot gets deployed, nobody tracks how many leads it captured, and six months later the business cannot decide whether it was worth the spend. We see this pattern in nearly every audit.

Our implementations invert the sequence. Before the engineer writes a line of configuration we define what the system is supposed to produce, how we will measure it, and what the baseline is. That baseline is captured during the audit. The measurement infrastructure deploys alongside the system, not after it.

The monthly report shows the measurement against the baseline, in dollars and hours. Not impressions. Not engagement.

What this means for you. At any point you can ask whether a specific system is working, and we will answer with data rather than hedging. If something is not producing the return we projected, we fix it or cut it. The measurement is the contract.

04

Tested before it ships

Evaluation suite

Graded against your real cases, not a demo script.

The demo always works. The demo is one example, run once, by the person who built it. What decides whether a system survives is how it behaves on the awkward cases your business actually produces, including the ones your own team gets wrong.

So before a system touches a customer we build an evaluation suite out of your own history: real enquiries, real tickets, real jobs. We write down what a correct response looks like for each one, then grade the system against them and fix what fails. You see the grade before launch, on cases we agreed in advance were the ones that mattered.

The suite is the part that lasts, and it is the reason we build it rather than testing by hand. A model gets updated. Someone edits a prompt. A connected tool changes the shape of what it returns. Any one of those can bring back a failure you already paid to fix, and none of them announce themselves. So we re-run the suite on every change we make and on every change made to us, which is a thing we do rather than a result we can promise.

What this means for you. You are shown test results, not a demo. When we say a system handles a case, we can produce the case and the grade. When it fails one, you hear it from us during the build.

05

Scoped authority and guardrails

Written before build

Every agent has a written list of what it may never do alone.

An agent that drafts a reply is a different proposition from one that sends it, and one that can spend money is different again. In the systems we audit, that distinction is almost never written down anywhere. The permissions end up being whatever the integration happened to grant on the day someone connected it, which is a decision nobody actually made.

We write the scope down before the build starts, in plain language rather than a permissions screen: what the agent may read, what it may write, which actions require a person to approve them, and what the agent does when it is not confident. Irreversible actions stay with a person. Money movement stays with a person.

The uncertain case is the one that usually gets skipped, and it is the one that causes the damage. A system that guesses where it should have asked is how you get the wrong booking, the wrong refund, the wrong quote. We define that handoff explicitly and then test it as one of the cases in the suite from pillar four.

None of this makes a system incapable of error, and we will not tell you otherwise. What we can describe is what we do about it. Irreversible actions stay with a person. What the agent did is logged, so a mistake can be found rather than inferred. And we agree in advance which failures have to stop the system rather than be absorbed by it.

What this means for you. You can read the scope before anything is built, and you approve it. If you later want an agent to hold more authority than that, it is a decision you make deliberately rather than one you discover after the fact.

06

Co-managed operations

Monthly, ongoing

We do not hand over a dashboard and leave.

The most common failure mode in AI implementation is also the most avoidable: the agency builds the system, hands over documentation and disappears. Six weeks later something breaks, nobody knows how to fix it, and the whole investment collapses.

Co-managed means we stay. Every engagement includes a monthly optimisation call where we review performance against the KPIs from pillar three, identify what is working and adjust what is not. It includes a shared channel, Slack or email, where your account manager is reachable during business hours. It includes documented changes logged in writing, so you always know what was adjusted and why.

This is different from "managed service" in the traditional sense, where the vendor owns the system and the client becomes dependent. Co-managed means the client has full access to everything we build: credentials, documentation, configuration files. You can take it in-house at any time. We stay because it is worth staying, not because you cannot leave.

What this means for you. You get an embedded team without the lock-in of a dependency. You can end the engagement at any time. We structured the model so that ending the engagement does not end the systems. The client owns everything we build.

Reference

Glossary

AI implementation carries a lot of jargon. Some of it is useful. Here is what we mean when we use these terms.

21 terms

Account manager

The non-technical half of your embedded team. Owns the relationship, runs the monthly reviews, and translates between your business goals and the engineer's implementation decisions. Named, reachable, and the same person throughout the engagement.

AI stack audit

Our diagnostic process at the start of every engagement. A structured inventory of every AI tool, subscription, workflow and automation running in your business: what it costs, what it produces, and where the gaps are.

Claude

Anthropic's model family. Our default for high-stakes reasoning, writing, and customer-facing conversation where accuracy and consistency under edge cases matter more than raw speed.

Co-managed operations

Our operating model after initial deployment. We remain accountable for running and optimising the systems we built, alongside your team. Monthly optimisation calls, shared communication channels, documented changes. You retain full access and ownership at all times.

Custom AI agent

A purpose-built agent we construct when off-the-shelf tools cannot do what the client needs. Built on open-source frameworks, configured for specific workflows, and owned by the client.

Discovery call

The 60-minute initial conversation. No pitch deck. We ask about your current operations, your AI spend, your top three frustrations and the outcomes you would want. By the end we have either identified a fit or we have not. Either is an acceptable outcome.

Evaluation suite

The set of graded test cases we build from a client’s own history before a system launches: real enquiries and real jobs, each with a written definition of what a correct response looks like. Re-run whenever the model, a prompt or a connected tool changes, because any of those can revive a failure that was already fixed.

Forward-deployed engineer

The technical half of your embedded team. Builds, configures and maintains the AI systems inside your business. Borrowed from firms like Palantir, the term describes technical implementers who work inside client environments rather than from an agency's remote office.

Hermes

An open-source AI agent framework built by Nous Research. Specialises in persistent memory and self-improving skills, so the agent gets more capable the longer it runs. We use it for personal agent deployments where long-term adaptation matters.

Implementation gap

The space between a tool that technically works and a tool that is producing value inside a specific business. Most AI failures happen here, not because the technology is broken, but because nobody closed the gap between the demo and the daily workflow.

KPI

The measurable outcome every implementation is tied to. Hours saved per week, leads captured per month, average response time, revenue attributed. Every system we deploy has a defined KPI before deployment.

Manus

An autonomous task-execution framework for agents that need to operate independently across multi-step workflows. Strong for research-intensive and execution-heavy use cases.

MCP, the Model Context Protocol

An open standard for connecting AI models to external tools and data sources, developed by Anthropic and adopted across the agent ecosystem. We build integrations on MCP because it is portable. The client owns the connections, not us.

OpenClaw

An open-source AI agent framework that lets agents execute real actions on a computer: file operations, browser control, API calls, shell commands. We use it for implementations requiring heavy multi-system automation. Client-owned and self-hosted.

Optimisation call

The standing monthly meeting for every co-managed engagement, typically 30 minutes. We review performance against KPIs, flag what is working and what is not, discuss proposed changes, and document decisions. Without it, co-managed becomes set and forget.

Perplexity Computer

A research-first agent environment. We reach for it when live web intelligence, citation trails and recurring monitoring are the centre of the task rather than a side effect.

ROI dashboard

The client-facing report we produce every month showing what each deployed system produced in measurable terms. Hours reclaimed. Leads captured. Calls answered. Revenue attributed. Delivered as a PDF or a shared page. Not a live dashboard of vanity metrics.

ROI-first methodology

Our organising principle. Every recommendation, implementation and optimisation decision is justified by expected or demonstrated return. "It would be interesting to add AI here" is not a sufficient reason to build something.

Scoped authority

The written statement of what an agent may read, what it may write, which actions require a person to approve them, and what the agent does when it is not confident. Agreed before the build rather than inherited from whatever permissions an integration happened to grant.

Stack sprawl

The condition most businesses arrive in: multiple AI tools accumulated over time, paid monthly, rarely audited, often redundant. Expensive, and it produces the illusion of AI capability without the measurable outcome.

Tool-agnostic

Our approach to vendor selection. We have no preferred tools we are paid to recommend. We recommend whatever fits the use case, budget and existing stack. The tool decision follows the audit, not the other way round.

Start the conversation

This is how we work. The first call is where we find out if it fits.