Data Quality Auditor
SkillDatabases & dataAudit data quality across six DAMA DMBOK dimensions, Completeness, Accuracy, Consistency, Timeliness, Validity, Uniqueness. Produces a scored data quality report with issue log and remediation plan. Use when profiling datasets, onboarding new data sources, or building data quality gates.
Use Data Quality Auditor in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Data Quality Auditor and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Data Quality Auditor skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; Ahel provides instructions and does not run this skill.
No other account needed.
Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
What this skill tells your AI
The instructions your AI receives, as published by hoavdc/codexkit in skills/codexkit-data-quality-auditor/SKILL.md and read by Ahel’s review.
When to Use
- Before building analytics or dashboards on a new data source
- When data issues cause downstream report errors
- During data migration or system integration
- When establishing data quality monitoring rules
Procedure
Step 1 — Scope & Profiling
Identify the dataset and context:
- Table/file name, row count, column count
- Business purpose: what decisions does this data support?
- Data owner: who is accountable?
Profile the data:
- Column types (string, numeric, date, boolean, null)
- Null rates per column
- Distinct value counts
- Min/Max/Mean for numerics
- Date range coverage
Step 2 — Six-Dimension Assessment
Score each dimension 1–5:
| Dimension | Question | Score Method |
|---|---|---|
| Completeness | Are all required fields populated? | % non-null for required fields |
| Accuracy | Do values represent reality? | Sample validation against source |
| Consistency | Do related fields agree? | Cross-field rule checks |
| Timeliness | Is data current enough for its purpose? | Max staleness vs requirement |
| Validity | Do values conform to allowed ranges/formats? | Format + range validation |
| Uniqueness | Are there unwanted duplicates? | Duplicate rate on key columns |
Step 3 — Issue Log
For each issue found:
| # | Dimension | Column(s) | Issue Description | Severity | Records Affected | Example |
|---|---|---|---|---|---|---|
| 1 | Completeness | 12% null in required field | High | 1,200 | row 45: null |
Severity: Critical (blocks use) / High (degrades quality) / Medium (cosmetic) / Low (nice-to-fix)
Step 4 — Remediation Plan
For each Critical/High issue:
| Issue | Root Cause | Fix | Owner | Deadline |
|---|---|---|---|---|
| Missing emails | Optional in old form | Backfill from CRM | Data Eng | Sprint 4 |
Step 5 — Monitoring Rules
Define ongoing quality checks:
- Automated checks to run on each data load
- Alerting thresholds (if quality drops below X%, alert owner)
- Review cadence (weekly, monthly)
Inputs
| Input | Required | Format |
|---|---|---|
| Dataset or table | Yes | Name + access method |
| Business context | Yes | What this data is used for |
| Quality rules / expectations | Recommended | Any existing rules or SLAs |
| Sample size | Recommended | Full scan or % sample |
Output
## Data Quality Report — [Dataset Name]
### Summary
- **Dataset:** orders_2024
- **Rows:** 145,000 | **Columns:** 23
- **Overall Quality Score:** 3.8 / 5.0
### Dimension Scores
| Dimension | Score | Key Finding |
|-----------|-------|-------------|
| Completeness | 4/5 | 2 columns with >5% nulls |
| Accuracy | 3/5 | 3% of prices don't match catalog |
| Consistency | 4/5 | state vs zip mismatch in 1.2% |
| Timeliness | 5/5 | Data refreshed within SLA |
| Validity | 3/5 | 8% of emails fail format check |
| Uniqueness | 4/5 | 0.5% duplicate order IDs |
### Critical Issues
[Issue log table]
### Remediation Plan
[Fix plan with owners and deadlines]
### Monitoring Rules
[Automated checks to add]
Definition of Done
- All six dimensions assessed and scored
- Issue log with severity and affected records
- Remediation plan for Critical/High issues
- Monitoring rules defined for ongoing quality
Quality Criteria
- Every finding is tied to a specific evidence source (log, test, metric)
- Pass/fail criteria are binary and measurable — no subjective judgments
- Severity levels are assigned with clear thresholds
- Remediation steps are provided for all critical and high findings
Verification (4C)
| Check | Question |
|---|---|
| Correctness | Are all pass/fail criteria applied against the correct standard or rule? |
| Completeness | Were all required dimensions or checklist items evaluated? |
| Context-fit | Does the verification scope match the actual risk level of the deliverable? |
| Consequence | If this passed verification but had a hidden flaw, what is the worst-case impact? |
Edge Cases
- Incomplete data for full assessment — Document which checks were limited and flag for re-verification when data becomes available.
- Ambiguous pass/fail criteria — Request clarification from the standard owner before scoring. Mark as 'Needs Review'.
- Multiple overlapping standards — Identify the governing standard and note where others diverge.
Changelog
- v1.0.0 — Initial release
Signals
- GitHub stars
- 25
- Forks
- 13
- Last commit
- Oct 2026
Advanced
- Item type
- skill
- Key
codexkit-data-quality-auditor- Source
- github.com/hoavdc/codexkit
Related picks
Skill · larksuite
The pick for Markdownmarkdown-formatter
Skill · nvidia
The pick for Markdownanalytics
Skill · coreyhaines31
More in Databases & datasupabase
Skill · supabase
More in Databases & dataconnect
Skill · composiohq
More in Databases & dataazure-kusto
Skill · microsoft
More in Databases & data