Arize evaluator

SkillDatabases & data

Create, update, and run Arize LLM-as-judge evaluators and tasks for spans, traces, sessions, projects, datasets, and experiments. Use when the user mentions create evaluator, LLM judge, hallucination, faithfulness, correctness, relevance, run eval, score spans, score experiment, trigger-run, column mapping, continuous monitoring, or evaluator prompt improvement.

Use Arize evaluator in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add Arize evaluator and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the Arize evaluator skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Arize evaluatorStart free

What this skill tells your AI

The instructions your AI receives, as published by gbb-acelerators/awesome-harness-primitives in harness/claude-code/skills/arize-evaluator/SKILL.md and read by Ahel’s review.

Design evaluator prompts, create or update Arize evaluators, map task columns, run evaluations, and troubleshoot ax failures without fabricating results or reading local credential files.

When to invoke

  • "Create an Arize LLM judge evaluator."
  • "Run a hallucination eval on spans."
  • "Score this experiment for correctness."
  • "Fix evaluator column mapping."
  • "Set up continuous monitoring with trigger-run."

Prerequisites and context

Proceed directly with the needed ax command; do not check versions, environment variables, or profiles up front. If an ax command fails, react to the error.

SymptomResolution
command not found or version errorRead references/ax-setup.md.
401 Unauthorized or missing API keyRun ax profiles show. If the profile is missing or wrong, use references/ax-profiles.md; if the user lacks a key, direct them to https://app.arize.com/admin > API Keys.
Space unknownRun ax spaces list and select by name, or ask the user.
LLM provider call fails because OPENAI_API_KEY or ANTHROPIC_API_KEY is missingRun ax ai-integrations list --space SPACE to check platform-managed credentials. If none exist, ask for a key or use the arize-ai-provider-integration skill.

Security rule: never read .env files or search the filesystem for credentials. Use ax profiles for Arize credentials and ax ai-integrations for LLM provider keys. If credentials are unavailable through those channels, ask the user.

SPACE, --space, and ARIZE_SPACE accept either a space name such as my-workspace or a base64 space ID such as U3BhY2U6...; find values with ax spaces list.

Evaluator model

FieldMeaning
TemplateJudge prompt with {variable} placeholders such as {input}, {output}, {context}, or {conversation}.
Classification choicesAllowed labels such as factual / hallucinated, correct / incorrect, or pass / fail; each may carry a numeric score.
AI IntegrationStored LLM provider credentials used by the evaluator.
ModelJudge model such as gpt-4o or claude-sonnet-4-5.
Invocation paramsJSON model settings such as {"temperature": 0}.
Optimization directionmaximize when higher scores are better, minimize when lower scores are better.
Data granularityspan, trace, or session; most evaluators default to span.

Evaluators are versioned. Every prompt or model change creates a new immutable version, and the newest version is active.

Task model and granularity

Task fieldMeaning
EvaluatorsOne or more evaluators to run.
Column mappingsMaps template variables to span or run fields, such as input → attributes.input.value.
Query filterSQL-style expression such as span_kind = 'LLM' to choose spans or runs.
ContinuousProject task option that scores new spans as they arrive.
Sampling rateContinuous task fraction from 0 to 1.
GranularityWhat it evaluatesUse forResult column prefix
spanIndividual spansQ&A correctness, hallucination, relevanceeval.{name}.label, eval.{name}.score, eval.{name}.explanation
traceSpans grouped by context.trace_idAgent trajectory and full call-chain task correctnesstrace_eval.{name}.label, trace_eval.{name}.score, trace_eval.{name}.explanation
sessionTraces grouped by attributes.session.id and ordered by start_timeMulti-turn coherence, tone, and conversation qualitysession_eval.{name}.label, session_eval.{name}.score, session_eval.{name}.explanation

For trace granularity, values are grouped by context.trace_id and comma-joined, with each value truncated to 100K characters. For session granularity, trace-level grouping happens first, then traces are ordered by start_time and grouped by attributes.session.id; session-level values are capped at 100K characters total. At session granularity, {conversation} renders as a JSON array of {input, output} turns from attributes.input.value / attributes.llm.input_messages and attributes.output.value / attributes.llm.output_messages. At span or trace granularity, {conversation} is resolved like any other mapped variable.

Multi-evaluator tasks may contain different granularities. Runtime uses the highest granularity, session > trace > span, and splits into one child run per evaluator. Per-evaluator query_filter in the task evaluators JSON narrows included spans, such as only tool-call spans within a session.

Template design rules

RuleRequirement
Portable variablesUse {input}, {output}, and {context}, not project-specific names such as {attributes_input_value}. Wire actual paths in column_mappings.
Binary firstPrefer two labels such as hallucinated / factual because more labels increase ambiguity and lower inter-rater reliability.
Exact label outputPrompt the judge to respond with only one label string, and ensure labels exactly match --classification-choices by spelling and casing.
Low temperatureUse --invocation-params '{"temperature": 0}' for reproducible scoring.
Explanations during setupUse --include-explanations while debugging judge behavior.
Shell quotingPass templates in single quotes, for example --template 'Judge this: {input} → {output}'; double quotes can cause shell interpolation.
Classification choicesAlways set --classification-choices; omitting it can fail with "missing rails and classification choices."

Limits

Never fabricate evaluation results. If a task fails, is cancelled, or produces no scores, report the failure and explain what happened. Do not perform a manual evaluation, invent quality scores, estimate percentages, or present agent analysis as Arize evaluation output. Recommend fixing the issue and retrying, trying the Arize UI, verifying credentials with ax ai-integrations list, or contacting https://arize.com/support.

Progressive disclosure and bundled resources

  • references/evaluator-crud-workflows.md: GraphQL CRUD calls, project and experiment evaluator setup, trigger-run operations, task management, column mapping, continuous monitoring, and troubleshooting.
  • references/ax-setup.md: ax install and version remediation.
  • references/ax-profiles.md: profile creation and update workflow.

Related primitives

NameTypeUse it when
arize-ai-provider-integrationskillCreating, updating, or deleting LLM provider credentials.
arize-traceskillExporting spans to discover column paths and time ranges.
arize-experimentskillCreating experiments and exporting runs for experiment column mappings.
arize-datasetskillExporting dataset examples to find input fields when runs omit them.
arize-linkskillCreating deep links to evaluators and tasks in the Arize UI.

Output template

### Arize evaluator result

**Status:** created | updated | run complete | failed | blocked
**Space:** `<SPACE or ARIZE_SPACE>`
**Evaluator:** `<name/id/version>`
**Task:** `<task id or n/a>`
**Granularity:** span | trace | session

| Step | Command | Result |
| --- | --- | --- |
| <step> | `ax ...` | <output summary> |

### Column mappings
| Template variable | Data field |
| --- | --- |
| `{input}` | `<field path>` |
| `{output}` | `<field path>` |

### Evaluation results
- <actual Arize result, task status, or failure reason; never fabricated>

Quality gate

  • The task proceeded with the needed ax command before speculative prechecks.
  • SPACE / --space / ARIZE_SPACE was resolved by name or base64 ID.
  • Credentials were checked only through ax profiles or ax ai-integrations; no .env files were read.
  • Template labels exactly match --classification-choices.
  • --invocation-params '{"temperature": 0}' is used unless a different temperature is justified.
  • Column mappings connect every template variable to real span, trace, session, project, dataset, or experiment fields.
  • Results reported are actual Arize outputs; failures are not converted into manual scores.

References

Signals

GitHub stars
20
Last commit
Sep 2026

Ahel review

  • K1binfo
    installs-packages (in references/ax-setup.md)

Automated review, not a security audit. Ruleset v1+k2.

Advanced
Item type
skill
Key
arize-evaluator-gbb-acelerators
Source
github.com/gbb-acelerators/awesome-harness-primitives