Scoping and evaluation sprint
Discovery that decides whether you need an agent, then writes the design: tools, permission scope per call, approval thresholds, evaluation scenarios and a model recommendation. It ends in an itemised pilot scope, or a written case for a cheaper build.
Anyone with a multi-step task they suspect could run itself, and anyone quoted an agent elsewhere who wants a second reading. It does not fit a task you already know is one question or one fixed path.
- The task tested for agent shape: many steps, several tools, a varying path
- Tools, permission scope per call and approval thresholds agreed in writing
- Evaluation scenarios from the real job, traps included
- A model recommendation on fit, cost and data residency
- An itemised pilot scope, or a written case for a cheaper build
- The pilot build
- Model and API usage during scenario testing
- Turnaround
- 1 to 2 weeks
- Moves the number
- How many systems the agent would need to touch: one CRM lookup and six back-office systems are different sprints.