Data Quality Auditor

SkillDatabases & data

Audit data quality across six DAMA DMBOK dimensions, Completeness, Accuracy, Consistency, Timeliness, Validity, Uniqueness. Produces a scored data quality report with issue log and remediation plan. Use when profiling datasets, onboarding new data sources, or building data quality gates.

Use Data Quality Auditor in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add Data Quality Auditor and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the Data Quality Auditor skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Data Quality AuditorStart free

What this skill tells your AI

The instructions your AI receives, as published by hoavdc/codexkit in skills/codexkit-data-quality-auditor/SKILL.md and read by Ahel’s review.

When to Use

  • Before building analytics or dashboards on a new data source
  • When data issues cause downstream report errors
  • During data migration or system integration
  • When establishing data quality monitoring rules

Procedure

Step 1 — Scope & Profiling

Identify the dataset and context:

  • Table/file name, row count, column count
  • Business purpose: what decisions does this data support?
  • Data owner: who is accountable?

Profile the data:

  • Column types (string, numeric, date, boolean, null)
  • Null rates per column
  • Distinct value counts
  • Min/Max/Mean for numerics
  • Date range coverage

Step 2 — Six-Dimension Assessment

Score each dimension 1–5:

DimensionQuestionScore Method
CompletenessAre all required fields populated?% non-null for required fields
AccuracyDo values represent reality?Sample validation against source
ConsistencyDo related fields agree?Cross-field rule checks
TimelinessIs data current enough for its purpose?Max staleness vs requirement
ValidityDo values conform to allowed ranges/formats?Format + range validation
UniquenessAre there unwanted duplicates?Duplicate rate on key columns

Step 3 — Issue Log

For each issue found:

#DimensionColumn(s)Issue DescriptionSeverityRecords AffectedExample
1Completenessemail12% null in required fieldHigh1,200row 45: null

Severity: Critical (blocks use) / High (degrades quality) / Medium (cosmetic) / Low (nice-to-fix)

Step 4 — Remediation Plan

For each Critical/High issue:

IssueRoot CauseFixOwnerDeadline
Missing emailsOptional in old formBackfill from CRMData EngSprint 4

Step 5 — Monitoring Rules

Define ongoing quality checks:

  • Automated checks to run on each data load
  • Alerting thresholds (if quality drops below X%, alert owner)
  • Review cadence (weekly, monthly)

Inputs

InputRequiredFormat
Dataset or tableYesName + access method
Business contextYesWhat this data is used for
Quality rules / expectationsRecommendedAny existing rules or SLAs
Sample sizeRecommendedFull scan or % sample

Output

## Data Quality Report — [Dataset Name]

### Summary
- **Dataset:** orders_2024
- **Rows:** 145,000 | **Columns:** 23
- **Overall Quality Score:** 3.8 / 5.0

### Dimension Scores

| Dimension | Score | Key Finding |
|-----------|-------|-------------|
| Completeness | 4/5 | 2 columns with >5% nulls |
| Accuracy | 3/5 | 3% of prices don't match catalog |
| Consistency | 4/5 | state vs zip mismatch in 1.2% |
| Timeliness | 5/5 | Data refreshed within SLA |
| Validity | 3/5 | 8% of emails fail format check |
| Uniqueness | 4/5 | 0.5% duplicate order IDs |

### Critical Issues
[Issue log table]

### Remediation Plan
[Fix plan with owners and deadlines]

### Monitoring Rules
[Automated checks to add]

Definition of Done

  • All six dimensions assessed and scored
  • Issue log with severity and affected records
  • Remediation plan for Critical/High issues
  • Monitoring rules defined for ongoing quality

Quality Criteria

  • Every finding is tied to a specific evidence source (log, test, metric)
  • Pass/fail criteria are binary and measurable — no subjective judgments
  • Severity levels are assigned with clear thresholds
  • Remediation steps are provided for all critical and high findings

Verification (4C)

CheckQuestion
CorrectnessAre all pass/fail criteria applied against the correct standard or rule?
CompletenessWere all required dimensions or checklist items evaluated?
Context-fitDoes the verification scope match the actual risk level of the deliverable?
ConsequenceIf this passed verification but had a hidden flaw, what is the worst-case impact?

Edge Cases

  • Incomplete data for full assessment — Document which checks were limited and flag for re-verification when data becomes available.
  • Ambiguous pass/fail criteria — Request clarification from the standard owner before scoring. Mark as 'Needs Review'.
  • Multiple overlapping standards — Identify the governing standard and note where others diverge.

Changelog

  • v1.0.0 — Initial release

Signals

GitHub stars
25
Forks
13
Last commit
Oct 2026
Advanced
Item type
skill
Key
codexkit-data-quality-auditor
Source
github.com/hoavdc/codexkit