AgentScore
AgentScore by Latitude evaluates production AI agents across outcome, reliability, cost, speed, and safety using ingested traces.
AgentScore by Latitude evaluates production AI agents across outcome, reliability, cost, speed, and safety using ingested traces.
What the product does and how it is positioned
AgentScore is an evaluation feature within the Latitude platform that converts production traces into a unified evaluation score. It evaluates AI agents across five key dimensions: outcome, reliability, cost, speed, and safety.
Rather than relying on synthetic benchmarks, the tool requires real production evidence sent via OpenTelemetry or existing trace sources. To display a score, an agent must clear predefined coverage, traffic, and confidence thresholds across a minimum of 1,000 eligible sessions.
Source-supported ways to use the product
Teams ingest live session traces via OpenTelemetry to measure agent success across five distinct operational dimensions.
Engineers inspect failed production sessions to identify negative behavioral patterns and convert them into ongoing evaluation benchmarks.
Developers feed production failure evidence into coding agents via MCP to implement targeted code fixes and verify their impact post-release.
The documented workflow, where available
Send live agent execution traces into Latitude manually via OpenTelemetry or through an existing tracing setup.
Let the platform collect at least 1,000 eligible sessions across a 7, 14, 21, or 28-day window until traffic, coverage, and confidence gates pass.
Review the resulting score, identify the dimensions holding back agent performance, and drill into individual failed sessions.
Configure evaluations based on identified production failure cases to monitor agent behavior over subsequent sessions.
Provide trace evidence to coding agents using MCP, deploy the updated agent code, and track the impact on subsequent production sessions.
AgentScore relies on strict statistical criteria before reporting a score. An agent must accumulate a minimum of 1,000 eligible production sessions before a score can be calculated. The evidence must pass traffic, coverage, and confidence gates across all five evaluation dimensions.
Scoring takes place over the shortest whole-week duration (7, 14, 21, or 28 days) that fulfills eligibility requirements. Each calculated score is accompanied by a 95% confidence interval, supporting uncertainty is clearly indicated when evaluating performance.
Checks to run with your own material and workflow
What was checked and when
Answers based on the source-checked product record
AgentScore evaluates production sessions across five core dimensions: outcome, reliability, cost, speed, and safety.
Traces can be connected via OpenTelemetry, brought from existing trace setups, or configured using the Latitude AI skill prompt.
An agent must accumulate at least 1,000 eligible production sessions and satisfy traffic, coverage, and confidence gates across all five dimensions.
AgentScore evaluates sessions over the shortest eligible whole-week window of 7, 14, 21, or 28 days that meets requirements.
Developers can inspect failure sessions, convert them into evaluations, and pass evidence directly to coding agents using MCP.