MCP Server Evaluation Creator
SkillMediaCreate evaluation Q&A pairs to test MCP server quality. Use when validating how well an MCP server enables LLMs to accomplish real tasks, benchmarking tool design, or measuring server effectiveness through realistic question-answer evaluations.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the MCP Server Evaluation Creator skill
What this skill tells your AI
The instructions your AI receives, as published by jmrplens/gitlab-mcp-server in .github/skills/create-mcp-evaluation/SKILL.md and read by ahel’s review.
Create comprehensive evaluations to test whether LLMs can effectively use your MCP server to answer realistic, complex questions using only the tools provided.
Overview
The quality of an MCP server is measured by how well its implementations (input/output schemas, descriptions, functionality) enable LLMs with no other context to answer realistic and difficult questions.
Process
Step 1: Tool Inspection
- List all tools available in the MCP server
- Understand input/output schemas, descriptions, and annotations
- Identify read-only vs. write operations
- Note pagination capabilities and limits
Step 2: Content Exploration
- Use READ-ONLY tools to explore available data
- Identify stable data points that won't change over time
- Map relationships between resources (projects → issues → comments → users)
- Note interesting patterns, edge cases, and complex relationships
Step 3: Question Design
Create 10 evaluation questions following these requirements:
Core Rules
- Questions MUST be independent (no dependency on other answers)
- Questions MUST require ONLY read-only, non-destructive operations
- Questions MUST be realistic — tasks humans with LLM assistance would care about
- Each answer MUST be a single, verifiable value (string comparison)
- Answers MUST be stable (won't change over time)
Complexity Guidelines
- Require multiple tool calls (potentially dozens)
- Multi-hop: answer depends on chaining information from multiple queries
- Require deep exploration, not surface-level keyword search
- Use synonyms and paraphrases, not direct keywords from target content
- May require extensive pagination through results
- Should stress-test tool return values across data modalities
Answer Diversity Cover diverse answer types:
- Names (user, project, group)
- IDs (project ID, issue IID)
- URLs and paths
- Timestamps and dates (specify format in question)
- Counts and quantities
- Boolean (True/False)
- Status values
Step 4: Verification
For each question:
- Solve it yourself using only the MCP tools
- Verify the answer is correct and stable
- Confirm it requires multiple tool calls
- Ensure no write operations are needed
Step 5: Project Evaluator Execution
For this repository, the evaluation is test/e2e/modeleval: it boots a GitLab,
drives the real cmd/server binary over stdio and puts the corpus to a model,
so what is measured is the surface a client is served. Cases are typed
declarations in internal/testutil/modelcorpus (ids such as MS-001,
MF-001), each carrying its own answer key. See
the developer guide
for the full setting matrix.
A prompt may not name its own answer. The stimulus you write must not
contain the tool, action or parameter names of that case's key, or
TestContract_NoStimulusNamesItsOwnAnswer fails: a prompt that hands the model
the call measures copying rather than tool selection. Where the literal really
is the request, declare it in declarations.go with its reason.
Rehearse without spending anything. fake:perfect replays each case's key
through the whole pipe and calls no provider:
MODELEVAL_MODELS=fake:perfect make modeleval-ce
Put a couple of cases to a real model, which needs a key, consent and a ceiling:
MODELEVAL_SPEND=yes \
MODELEVAL_BUDGET_USD=5 \
MODELEVAL_CASES=MS-001,MF-001 \
MODELEVAL_MODELS='anthropic:claude-haiku-4-5-20251001' \
make modeleval-ce
Output Format
<evaluation>
<qa_pair>
<question>Find the GitLab project that contains a CI/CD pipeline with a job named "deploy-prod". What is the default branch of that project? Answer with the exact branch name.</question>
<answer>main</answer>
</qa_pair>
<qa_pair>
<question>In the project with ID 42, find the merge request that was merged on 2024-01-15. Who authored it? Answer with the username.</question>
<answer>jdoe</answer>
</qa_pair>
<!-- 8 more qa_pairs -->
</evaluation>
Example Question Patterns for GitLab MCP
- Cross-resource lookup: "Find the issue in project X that mentions Y. Who was assigned to it?"
- Multi-hop navigation: "Find the MR that fixed issue #N. What pipeline job failed first in that MR?"
- Aggregation: "How many open issues with label 'bug' exist across all projects in group Z?"
- Temporal reasoning: "What was the last commit to the develop branch of project X before 2024-06-01? Answer with the short SHA."
- Relationship tracing: "Find the user who created the most merge requests in project X during 2024. Answer with their username."
Quality Checklist
- 10 questions created
- All questions use only read-only operations
- All questions are independent
- All answers verified manually
- Answers are stable (won't change)
- Answers use direct string comparison
- Questions cover diverse tool usage (projects, issues, MRs, pipelines, users)
- Questions require multi-step reasoning
- No questions solvable with a single tool call
- Output format matches XML schema
Signals
- GitHub stars
- 39
- Forks
- 5
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
create-mcp-evaluation- Source
- github.com/jmrplens/gitlab-mcp-server