Skip to main content
Helicone OSS LLM Observability

Eval Scores

2 min read

When running evaluation frameworks to measure model performance, you need visibility into how well your AI applications are performing across different metrics. Scores let you report evaluation results from any framework to Helicone, providing centralized observability for accuracy, hallucination rates, helpfulness, and custom metrics.

Note

Helicone doesn't run evaluations for you - we're not an evaluation framework. Instead, we provide a centralized location to report and analyze evaluation results from any framework (like RAGAS, LangSmith, or custom evaluations), giving you unified observability across all your evaluation metrics.

Why use Scores#

  • Centralize evaluation results: Report scores from any evaluation framework for unified monitoring and analysis
  • Track model performance over time: Visualize how accuracy, hallucination rates, and other metrics evolve
  • Compare experiments side-by-side: Evaluate different prompts, models, or configurations with consistent metrics

Quick Start#

Run your evaluation

Use your evaluation framework or custom logic to assess model responses and generate scores (integers or booleans) for metrics like accuracy, helpfulness, or safety.

Report scores to Helicone

Send evaluation results using the Helicone API:

Alternative: Add scores via dashboard

You can also add scores directly in the Helicone dashboard on the request details page. This is useful for manual evaluation or quick testing.

View score analytics

Analyze evaluation results in the Helicone dashboard to track performance trends, compare experiments, and identify areas for improvement.

Warning

Scores are processed with a 10 minute delay by default for analytics aggregation.

API Format#

Request Structure#

The scores API expects this exact format:

FieldTypeDescriptionRequiredExample
scoresobjectKey-value pairs of evaluation metrics✅ Yes{"accuracy": 92}

Score Values#

TypeDescriptionExample
integerNumeric scores (no decimals)92, 85, 0
booleanPass/fail or true/false metricstrue, false
Warning

Float values like 0.92 are rejected. Convert to integers: 0.9292

Use Cases#

Evaluate retrieval-augmented generation for accuracy and hallucination: