OpenJudge Skill

SkillAI & models

openjudge is a skill that lets your AI build and run custom evaluation pipelines that grade outputs from language models. Once added, your AI can judge responses with graders based on a language model, a function, or an agent, and combine the results into final scores. This lets it check output quality in batches instead of reviewing each answer by hand.

Available today. Use it from your connected AI after setup.

After adding it, tell your agent what you want to evaluate and how quality should be judged, then ask it to build an evaluation pipeline with suitable graders. Run a batch evaluation to get grades for your outputs.

Then ask your AI: use the OpenJudge Skill skill

What your AI can do with it

  • Build custom evaluation pipelines for grading LLM outputs
  • Select and configure LLM-based, function-based, or agentic graders
  • Run batch evaluations across many outputs
  • Combine scores from multiple graders with aggregators
  • Apply voting or averaging strategies to settle final grades
  • Auto-generate graders

What this skill tells your AI

The instructions your AI receives, as published by agentscope-ai/openjudge in skills/openjudge/SKILL.md and read by ahel’s review.

Build evaluation pipelines for LLM applications using the openjudge library.

When to Use This Skill

  • User wants to evaluate LLM output quality (correctness, relevance, hallucination, etc.)
  • User wants to compare two or more models and rank them
  • User wants to design a scoring rubric and automate evaluation
  • User wants to analyze evaluation results statistically
  • User wants to build a reward model or quality filter

Sub-documents — Read When Relevant

TopicFileRead when…
Grader selection & configurationgraders.mdUser needs to pick or configure an evaluator
Batch evaluation pipelinepipeline.mdUser needs to run evaluation over a dataset
Auto-generate graders from datagenerator.mdNo rubric yet; generate from labeled examples
Analyze & compare resultsanalyzer.mdUser wants win rates, statistics, or metrics

Read the relevant sub-document before writing any code.

Install

pip install py-openjudge

Architecture Overview

Dataset (List[dict])
    │
    ▼
GradingRunner                    ← orchestrates everything
    │
    ├─► Grader A ──► EvaluationStrategy ──► _aevaluate() ──► GraderScore / GraderRank
    ├─► Grader B ──► EvaluationStrategy ──► _aevaluate() ──► GraderScore / GraderRank
    └─► Grader C ...
    │
    ├─► Aggregator (optional)    ← combine multiple grader scores into one
    │
    └─► RunnerResult             ← {grader_name: [GraderScore, ...]}
            │
            ▼
        Analyzer                 ← statistics, win rates, validation metrics

5-Minute Quick Start

Evaluate responses for correctness using a built-in grader:

import asyncio
from openjudge.models.openai_chat_model import OpenAIChatModel
from openjudge.graders.common.correctness import CorrectnessGrader
from openjudge.runner.grading_runner import GradingRunner

# 1. Configure the judge model (OpenAI-compatible endpoint)
model = OpenAIChatModel(
    model="qwen-plus",
    api_key="sk-xxx",
    base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
)

# 2. Instantiate a grader
grader = CorrectnessGrader(model=model)

# 3. Prepare dataset
dataset = [
    {
        "query": "What is the capital of France?",
        "response": "Paris is the capital of France.",
        "reference_response": "Paris.",
    },
    {
        "query": "What is 2 + 2?",
        "response": "The answer is five.",
        "reference_response": "4.",
    },
]

# 4. Run evaluation
async def main():
    runner = GradingRunner(
        grader_configs={"correctness": grader},
        max_concurrency=8,
    )
    results = await runner.arun(dataset)

    for i, result in enumerate(results["correctness"]):
        print(f"[{i}] score={result.score}  reason={result.reason}")

asyncio.run(main())

Expected output:

[0] score=5  reason=The response accurately states Paris as capital...
[1] score=1  reason=The response gives the wrong answer (five vs 4)...

Key Data Types

TypeDescription
GraderScorePointwise result: .score (float), .reason (str), .metadata (dict)
GraderRankListwise result: .rank (List[int]), .reason (str), .metadata (dict)
GraderErrorError during evaluation: .error (str), .reason (str)
RunnerResultDict[str, List[GraderResult]] — keyed by grader name

Result Handling Pattern

from openjudge.graders.schema import GraderScore, GraderRank, GraderError

for grader_name, grader_results in results.items():
    for i, result in enumerate(grader_results):
        if isinstance(result, GraderScore):
            print(f"{grader_name}[{i}]: score={result.score}")
        elif isinstance(result, GraderRank):
            print(f"{grader_name}[{i}]: rank={result.rank}")
        elif isinstance(result, GraderError):
            print(f"{grader_name}[{i}]: ERROR — {result.error}")

Model Configuration

All LLM-based graders accept either a BaseChatModel instance or a dict config:

# Option A: instance
from openjudge.models.openai_chat_model import OpenAIChatModel
model = OpenAIChatModel(model="gpt-4o", api_key="sk-...")

# Option B: dict (auto-creates OpenAIChatModel)
model_cfg = {"model": "gpt-4o", "api_key": "sk-..."}
grader = CorrectnessGrader(model=model_cfg)

# OpenAI-compatible endpoints (DashScope / local / etc.)
model = OpenAIChatModel(
    model="qwen-plus",
    api_key="sk-xxx",
    base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
)

Signals

GitHub stars
830
Forks
69
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
openjudge
Source
github.com/agentscope-ai/openjudge
OpenJudge Skill: Skill · ahel