assess-construct-validity

SkillDev tools

Lets your agent check whether a benchmark actually measures the capability it claims to measure.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the assess-construct-validity skill

About this skill

Assess whether a benchmark measures its claimed capability using content, convergent, discriminant, and confound analysis.

What this skill tells your AI

The instructions your AI receives, as published by yogsoth-ai/de-anthropocentric-research-engine in skills/assess-construct-validity/SKILL.md and read by ahel’s review.

Purpose

Assess whether a benchmark or evaluation construct measures the claimed capability rather than content cues, confounds, or unrelated skill.

Input contract

required: [construct_claim, benchmark_specification, evaluation_records]
optional: [content_analysis, convergent_measures, discriminant_measures, confound_hypotheses]
constraints: [each validity judgment requires an observable indicator, comparison basis, and provenance]

Procedure

  1. State the target construct and map benchmark tasks, labels, and metrics to its intended components.
  2. Check content coverage and plausible construct-irrelevant cues against the benchmark specification.
  3. Compare convergent and discriminant evidence where available, preserving missing comparisons.
  4. Test confound hypotheses with controlled contrasts or artifact probes and record residual uncertainty.

If construct validity depends on whether the operationalization covers the intended domain, consider map-coverage-space as the next tactic.

Output contract

produces: [construct_map, content_validity_assessment, convergent_discriminant_evidence, confound_report, validity_judgment]
delta_fields: [findings, evidence_updates, uncertainties, decisions, open_questions]

Quality gates

  • Every validity claim cites a task, measure, contrast, or artifact observation.
  • Content coverage, convergence, discrimination, and confounds are reported separately.
  • A missing diagnostic is marked unresolved rather than treated as evidence of validity.

Failure and counterexamples

Do not infer construct validity from a high score, face validity, or agreement with another measure that shares the same confound.

Provenance map

  • resolved: construct-validity-assessment

Signals

GitHub stars
501
Forks
41
Last commit
Sep 2026
Advanced
Catalog kind
skill
Key
assess-construct-validity
Source
github.com/yogsoth-ai/de-anthropocentric-research-engine