AgentScore  favicon

AgentScore

AgentScore by Latitude evaluates production AI agents across outcome, reliability, cost, speed, and safety using ingested traces.

Code & ITCalculates an aggregate score based on…Inspects specific production trace spansSurfaces recurring failure patterns and…Converts observed production failures…
AgentScore  product interface screenshot
Estimated monthly visits
101K
Data period:
Listed on AIToolly

What Is AgentScore ? Product Overview

What the product does and how it is positioned

AgentScore is an evaluation feature within the Latitude platform that converts production traces into a unified evaluation score. It evaluates AI agents across five key dimensions: outcome, reliability, cost, speed, and safety.

Rather than relying on synthetic benchmarks, the tool requires real production evidence sent via OpenTelemetry or existing trace sources. To display a score, an agent must clear predefined coverage, traffic, and confidence thresholds across a minimum of 1,000 eligible sessions.

What Can You Use AgentScore For?

Source-supported ways to use the product

Evaluating production agent quality

Teams ingest live session traces via OpenTelemetry to measure agent success across five distinct operational dimensions.

Failure analysis and eval creation

Engineers inspect failed production sessions to identify negative behavioral patterns and convert them into ongoing evaluation benchmarks.

Guiding fixes with coding agents

Developers feed production failure evidence into coding agents via MCP to implement targeted code fixes and verify their impact post-release.

How to Use AgentScore

The documented workflow, where available

  1. 1

    Connect production traces

    Send live agent execution traces into Latitude manually via OpenTelemetry or through an existing tracing setup.

  2. 2

    Accumulate evidence and meet gates

    Let the platform collect at least 1,000 eligible sessions across a 7, 14, 21, or 28-day window until traffic, coverage, and confidence gates pass.

  3. 3

    Inspect failure patterns

    Review the resulting score, identify the dimensions holding back agent performance, and drill into individual failed sessions.

  4. 4

    Turn failures into evals

    Configure evaluations based on identified production failure cases to monitor agent behavior over subsequent sessions.

  5. 5

    Ship fixes and verify

    Provide trace evidence to coding agents using MCP, deploy the updated agent code, and track the impact on subsequent production sessions.

Scoring Methodology and Eligibility Gates

AgentScore relies on strict statistical criteria before reporting a score. An agent must accumulate a minimum of 1,000 eligible production sessions before a score can be calculated. The evidence must pass traffic, coverage, and confidence gates across all five evaluation dimensions.

Scoring takes place over the shortest whole-week duration (7, 14, 21, or 28 days) that fulfills eligibility requirements. Each calculated score is accompanied by a 95% confidence interval, supporting uncertainty is clearly indicated when evaluating performance.

  • Covers five distinct dimensions: outcome, reliability, cost, speed, and safety.
  • Requires at least 1,000 eligible production sessions to satisfy data gates.
  • Calculates metrics across whole-week windows of 7, 14, 21, or 28 days.
  • Provides a 95% confidence interval alongside the calculated score.
  • Suppresses score output when incoming evidence does not satisfy gating requirements.

What to Test Before Choosing AgentScore

Checks to run with your own material and workflow

  • Confirm that your application architecture supports sending OpenTelemetry traces or integrating with existing trace data.
  • Verify that your agent generates sufficient production volume to meet the minimum eligibility threshold of 1,000 sessions.
  • Check whether your development workflow can leverage Model Context Protocol (MCP) to pass trace evidence to coding agents.
  • Review the requirements for coverage and confidence gates across all five scoring dimensions before expecting a score.

AgentScore Sources and Last Checked

What was checked and when

Last checked
Category
Code & IT

AgentScore Frequently Asked Questions

Answers based on the source-checked product record

What dimensions are measured by AgentScore?

AgentScore evaluates production sessions across five core dimensions: outcome, reliability, cost, speed, and safety.

How are production traces ingested into the platform?

Traces can be connected via OpenTelemetry, brought from existing trace setups, or configured using the Latitude AI skill prompt.

What are the minimum eligibility requirements to receive a score?

An agent must accumulate at least 1,000 eligible production sessions and satisfy traffic, coverage, and confidence gates across all five dimensions.

What time windows does AgentScore use for evaluation?

AgentScore evaluates sessions over the shortest eligible whole-week window of 7, 14, 21, or 28 days that meets requirements.

How can developers use AgentScore evidence to resolve issues?

Developers can inspect failure sessions, convert them into evaluations, and pass evidence directly to coding agents using MCP.

Explore other recently added tools in the same category.