SideKind
← All field notes

How to build an AI agent: the decisions only you can make

·
A build timeline in three phases: design, build, and rollout. Five numbered decision cards sit above the design phase: the one job, what you hand over, who signs off, team time, and the signed scope. The first card connects to the start of design and the second to a point early in design; the last three join one line that meets the timeline at the gate between design and build. Below the timeline, an owner lane under design says all five are due by the end of design. A dark builder lane under build lists the model, the tools, and how many agents. Under rollout, three identical labels list the stages: shadow, supervised, and autonomous. A footer reads: your five decisions come before the build, and the builder's choices sit inside it.

Five decisions in an AI agent build belong to the business that pays for it. Make them late and the build either waits for you or guesses.

This is for the owner or ops lead who has already decided to build. Maybe you hired a studio, maybe a freelancer, maybe your own engineer. Whether the loop deserves an agent at all is a separate question with its own diagnostic.

The engineering choices belong to whoever writes the code, and you judge those by what the agent produces. The five below are different. Nobody else can make them for you.

The five decisions

Each of the five decisions has a point in the build where it falls due; miss that point and someone else makes the call by default.

# The decision Due by If nobody makes it
1 What is the agent's one job? Before design starts The builder picks the job that is easiest to demo
2 What do you hand over? The first week of design The agent learns your process from guesses
3 Who on your side signs off? Before you sign the scope Questions wait days for an answer
4 How much of your team's time goes in? Before you sign the scope Reviews slip and the rollout stalls
5 What does the scope say? The end of design You approve a build you cannot check

1. What do you need to decide before building an AI agent?

Decide the agent's job in one paragraph: where the loop starts, where it stops, and what it hands back to a person. Everything else in the build is sized from that paragraph.

Write it about one loop. "Handle our inbound" is a department. "Draft a first reply to every inbound form fill and flag the ones that mention pricing" is a job.

Gartner predicts that more than 40% of agentic AI projects will be canceled by the end of 2027. Unclear business value is one of the three causes Gartner names. A job paragraph is the cheapest defense you have against that one.

Anthropic's engineering guidance says success with language models "isn't about building the most sophisticated system". Hold your builder to that line when the paragraph starts to grow.

2. What does your builder need from you before the build starts?

Your builder needs the procedure as your team runs it today, plus twenty recent real cases with the exceptions left in. Then they need working access to every system the job touches.

The procedure matters more than it looks. OpenAI's guide to building agents tells builders to base an agent's instructions on "existing operating procedures, support scripts, or policy documents". If yours are not written down, your builder has nothing to start from.

A 2018 survey of 1,001 U.S. workers, run by the video vendor Panopto with YouGov, found that 42 percent of institutional knowledge is unique to the individual. Panopto sells knowledge-sharing software, so discount accordingly.

Hand all of it over in the first week of design. If the procedure still lives in one person's head, map the workflow before you automate it. Deciding what the agent may change is its own decision, with its own field note on permissions.

3. Who on your side signs off during an AI agent build?

One named person signs off for the business. Call them your owner: often the one who does the work today, or their manager. They do not need to be an engineer.

During the build your owner has three duties. They sign the test sentence that says the agent is working. When the builder asks how an edge case should go, they answer within a day. At the end, they accept the build or send it back.

The test sentence is the one that matters most. It has to name a number, a population, and a place to measure it, and it has to exist before the build starts. Writing one well is its own piece of work.

A committee cannot sign off on an agent. When three people share the job, every design question becomes a meeting. Your owner often keeps reading the audit trail after launch, which is a role of its own.

4. How much of your team's time does building an AI agent take?

Plan for most of your team's time to land in two places: the design weeks and the first weeks of the rollout. The build in between needs a few check-ins.

On a SideKind engagement, design and scope take five to ten business days. Budget a few hours of your owner's time across those days to walk through the cases and read a draft of the scope.

A SideKind build goes from signed scope to live in 3 to 4 weeks, with a phased rollout: shadow mode first, then supervised, then autonomous. Budget more of your owner's time for the rollout than for design. In shadow mode, someone on your side should read what the agent would have done, run by run. Who decides when each stage ends has its own field note.

If your owner cannot find that review time, move the start date. A rollout nobody reviews is a rollout nobody can defend later.

5. What should an AI agent scope document include?

A scope document should hold at least four things: the one-paragraph job, every system the agent touches, every action it may take, and the signed test sentence.

Read the action list line by line. Each entry is something the agent does in your name. OpenAI's guide says actions that are "sensitive, irreversible, or have high stakes" should trigger human oversight. Mark those in the scope, before anyone writes code.

The scope should also say what happens when the agent is stuck: who receives the handoff, and what arrives with it.

Last, check the commercial terms against the scope. A fixed price only protects you if the scope is specific enough to price. Why a build is scoped that way has its own field note.

What stays with the builder

Your job is to hold the test.

Which model, which tools, whether to split the work across several agents: those are your builder's calls. You can ask about them. You should not have to referee them.

Anthropic's own team reported that on one coding agent they "spent more time optimizing our tools than the overall prompt". That work is real, and it is invisible from where you sit. Judge it by the output.

Three questions let you do that without reading any code.

  • Which record proves the agent did it? A log line that says "ok" can mean only that nothing crashed.
  • What does one run cost? The build is a fixed cost and running it is a recurring one. The arithmetic for both sits with the cost field note.
  • What happens when a tool fails? You want a named fallback, and the failure should be visible to your owner.

Common questions

Should we build it in-house or hire someone?

Either can work, and the five decisions stay yours in both cases. An in-house engineer still needs a job paragraph and a signed test, plus your owner reading the trail. What changes is who carries the estimation risk and how fast the build starts. Building in-house against hiring a team gets its own comparison.

Does our data need cleaning up before the build?

Only the data the agent's one job reads needs attention before the build. Check those fields against the twenty recent cases you are handing over, and note where they are missing or often wrong. Put the known gaps in the scope. A builder can plan around a gap you named. A gap found in shadow mode costs rollout time instead.

How do we avoid being locked in to our builder?

Settle what you own in the scope, before the build starts, so the agent can leave with you if the relationship ends. Asking for it after launch puts you in the weaker position. On a SideKind engagement you can cancel any month and take it in-house any time. What a clean handoff contains has its own field note.

What the call covers

The diagnostic call runs thirty minutes, free, with no pitch. You bring the loop. You leave with a blunt yes, no, or not-yet, and if it's a yes, a one-page write-up of what we'd scope. The whole process is written out here. Book the call here.

Write the job paragraph before the call; it makes the thirty minutes count.

Want to chat about this?

We love hearing from people who've thought hard about the same problems.