Skip to main content
After observing an application in production, the next step is annotating and curating that data to build evaluation datasets. This process turns raw production logs into high-quality test cases that drive systematic improvement. Use Cases
Capture thumbs up/down ratings, custom scores, or categorical labels on AI responses. Build a feedback loop that surfaces low-quality generations for review.
Flag responses with specific defects (hallucination, off-topic, inappropriate content) using structured annotation keys shared across the team.
Annotate Traces with corrections and quality labels, then export curated subsets as training datasets for future experiments.
Route Traces to Annotation Queues for systematic expert review. Combine with Trace Automations to automatically surface Traces that meet specific criteria.
Concepts Three concepts work together to form the annotations system:
  • Annotations: the feedback categories configured for a project, such as a quality rating or a defect tag
  • Annotation Queues: organized workflows for reviewing Traces in bulk via AI Studio
  • Annotations API: the API and SDK for applying feedback values to a Trace or span programmatically

Annotations

Define annotation schemas: keys, value types, and validation rules. Available on chat completion and responses spans once created.

Annotation Queues

Organize annotation review workflows. Filter and present relevant Traces for review in bulk.

Annotations API

Apply structured human feedback to Traces and spans programmatically via the API and SDK.

Create Annotations

Each annotation must be defined in the project before it can be used. The definition sets the key, title, and value type; applying an annotation to a Trace requires matching one of these definitions.
To create an annotation, head to Optimization > Annotations and press the button. Annotations can also be created directly from an Annotation Queue.
Create annotation form with Key, Title, Description fields and a Type selector showing Categorical, Range, and Text options.

Customizing an Annotation.

Each annotation uses one of three value types:
  • Categorical: button options with custom labels, such as good/bad or saved/deleted
  • Range: a custom scoring slider, for example a scale from 0 to 100
  • Open field: free-form text input for detailed comments
Once created, an annotation is available on all chat completion spans and responses spans in the project. No additional configuration or filtering required.
Deleting an annotation removes it from any Annotation Queues and Experiments that use it, so it no longer appears as a review option there. Annotations already recorded on a Trace are preserved: every annotated data point remains stored and queryable.

Common Annotations Legacy

Rate the overall quality of AI responses:
Identify specific issues with AI responses:
Multiple defects can be selected for one response using an annotation configured with an array value.

Use Annotations

Annotations can be applied wherever a Trace or span is reviewed:
  • Directly on a Trace or Log: open a single Trace or Log in the Traces or Logs view and use the Annotations panel.
  • In an Annotation Queue: review a curated set of Traces in bulk. Fill a queue with Trace Automations or by manually adding individual Traces or Logs.
  • Programmatically: apply feedback through the API and SDK using the API & SDK tab below.
  • In an Experiment: apply annotations while reviewing experiment outputs.
Every annotation applied in an Annotation Queue is written back to its originating Trace. Because the values live on the Trace, they can be queried with the Orq MCP and used to run analysis across reviewed data.
The annotation capabilities differ between Logs and Traces. Logs support human feedback and text corrections to the AI response. Traces support human feedback and correcting an evaluator result.
Navigate to the Traces view and select a single trace. The Annotations panel will be displayed, allowing you to apply human feedback to the AI response.
Trace detail panel for a claude-sonnet chat-completion showing Evaluations section with Defects, Interactions, and Rating feedback options including good/bad thumbs.

The Annotations panel in Traces lets you apply human feedback.

Annotations in Experiments

Annotations can also be applied outside of Annotation Queues, while reviewing the outputs of an Experiment. In the experiment review screen, the annotations defined for the project appear alongside Evaluator scores, so outputs can be annotated manually as part of an evaluation run.
Experiment review screen showing Response 1 of 20 for product-orchestrator-A, a left panel with Inputs, Expected, and Metrics, a center panel with the System instructions, User input, and Assistant output including function calls, and a right panel with an Annotations comment field and good/bad rating above an Evaluators section listing a json_check evaluator marked No.

The Annotations panel in the experiment review screen, with annotations shown above the Evaluator scores.

Correcting an Evaluator result is not available in the experiment review screen. Annotations and Evaluator scores still display as described above.