Measure an AI agent as a complete run, not a session

By Pascal Bouman··3 min read
Dashboard concept for measuring AI agent runs rather than ordinary sessions.

The product choice: manage by the run

Choose the run, not the session, as the primary unit for AI agent product analytics. A session tells you that someone used the product; it does not say which task the agent performed, which intermediate steps were needed, or why the outcome was or was not useful. NIST describes agents as systems that plan multi-step tasks and can autonomously perform actions, including using tools and searching databases. That source therefore supports the need to make execution visible, but does not prescribe a universal analytics schema. Define a run as one bounded goal with a start condition and an end status. At a minimum, capture: the goal and relevant input, selected steps, tool calls, escalation to a human, error or deviation, recovery action, final result, and user acceptance. This lets you distinguish between a technically completed task and a result the user actually adopts. That distinction is a product choice, not an effect on quality or revenue proven by the sources.

Make deviations and recovery actionable

Use run data first to find friction in one specific workflow. NIST reports that one evaluated model version had problems with tool calling and output formatting; researchers adjusted prompts, among other measures, and added simple mechanisms to support recovery from errors. This is evidence that such issues and recovery mechanisms occurred in that evaluation, not that the same solution works for every agent. Therefore, for every error, record both the affected step type and the recovery path and final acceptance. For applications that are considered high-risk within the scope of the AI Act, the European Commission lists activity logging for result traceability and detailed documentation as obligations before market introduction. This makes a run log especially relevant in that context; it does not mean that every AI agent falls under these rules. Link the measurement outcome to a decision: simplify a step, add a checkpoint, restrict a tool, or have a type of task escalated.

Flow of an AI agent run with steps and escalation points.

A run card for the next evaluation

Do not start with a large data model. Choose one common task and fill in the same fields for each run: run ID, goal, input category, steps, tools, deviation, recovery, human review, end status, acceptance, and owner of the follow-up. Then compare only runs with the same task goal. This reveals which step repeatedly requires recovery or escalation, without using session duration as a substitute for quality. Practical tool: Use a run card with the goal, owner, steps, tool use, escalation, error, recovery action, acceptance, and follow-up decision for one bounded agent workflow. Review rejected or recovered runs weekly and select one change to test afterward in the same run category. Limitation: The available passages support the relevance of agents' multi-step workflows, tool use, visibility, recovery issues and — in the high-risk context mentioned — traceability obligations. They provide no acceptance thresholds, no complete measurement schema, and no evidence that run measurement by itself causes better performance.

Your personal AI research team

Developments move too fast to keep up with everything yourself.

You need a research team that tracks changes, checks sources and decides what matters for your work.

Choose what you want to follow and receive only the updates that matter to you.

Updates tailored to your interests
Researched by specialist agents
Relevant insights, not daily noise

What do you want to follow?

You receive a confirmation email first and only join after clicking it. See the privacy policy.

Latest articles

Recent knowledge base articles selected for this page.