Masking PII Data

SkillDev tools

Protect personally identifiable information in data pipelines, classifying PII, choosing masking vs tokenization vs hashing vs encryption, dynamic data masking and column-level access control, and handling deletion/right-to-be-forgotten. Use when handling sensitive data, masking or anonymizing PII, meeting GDPR/CCPA/HIPAA requirements, or restricting column access in a warehouse.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Masking PII Data skill

What this skill tells your AI

The instructions your AI receives, as published by unknown-333/awesome-data-engineering-skills in skills/masking-pii-data/SKILL.md and read by ahel’s review.

When to use

  • A pipeline or table contains personal/sensitive data (names, emails, SSNs, payment, health).
  • Choosing how to de-identify data for analytics or lower environments.
  • Restricting who can see raw sensitive columns.
  • Handling deletion / right-to-be-forgotten requests.
  • Do NOT use for general access control unrelated to sensitive data.

Choose the right technique

TechniqueReversibleKeeps analytics utilityUse for
Masking / redactionNoLowDisplay, lower envs (j***@x.com)
Hashing (salted)NoJoin/match onlyPseudonymous keys, dedup
TokenizationYes (via vault)Referential joinsReversible pseudonymization
EncryptionYes (with key)None until decryptAt-rest protection, restricted fields

Workflow

- [ ] Classify columns: what is PII/sensitive and its risk level
- [ ] Pick technique per column by whether you need reversibility/joins
- [ ] Apply as early as possible (mask on ingest for lower environments)
- [ ] Enforce column-level access / dynamic masking for raw data
- [ ] Support deletion: know every location a subject's data lives
  1. Classify first. You can't protect what you haven't identified; tag columns by sensitivity. Pair with designing-data-contracts to declare PII fields.
  2. Pick per column. Need to join across systems but not reverse? Salted hash. Need to recover the value later? Tokenization/encryption. Just hide it? Mask.
  3. Apply early. Mask/tokenize before data reaches analysts or dev/test environments; never copy raw PII into lower environments.
  4. Access control. Use warehouse dynamic data masking and column-level grants so only authorized roles see raw values.
  5. Deletion. Track where each subject's data lives (lineage helps) so erasure requests are complete.

Patterns

Snowflake dynamic masking policy (unmask only for a privileged role):

CREATE MASKING POLICY email_mask AS (val string) RETURNS string ->
  CASE WHEN CURRENT_ROLE() IN ('PII_READER') THEN val
       ELSE REGEXP_REPLACE(val, '^[^@]+', '***') END;
ALTER TABLE customers MODIFY COLUMN email SET MASKING POLICY email_mask;

Salted hash for pseudonymous joins — hash with a secret salt so the same person matches across tables without exposing the raw identifier; keep the salt in a secrets manager.

Tokenization — replace the value with a token and store the mapping in a restricted vault; analytics use the token, authorized systems detokenize.

Common pitfalls

  • Copying raw PII to dev/test — the most common leak; mask on the way down.
  • Unsalted hashes — vulnerable to rainbow tables and re-identification; always salt.
  • Masking at display only while storing raw everywhere — breach still exposes data; protect at rest and restrict access.
  • Forgetting free-text/logs — PII hides in comments, logs, and JSON blobs, not just typed columns.
  • No deletion plan — right-to-be-forgotten fails if you can't locate all copies; use lineage.
  • Reversible where you meant irreversible — don't tokenize when the requirement is true anonymization.

Signals

GitHub stars
21
Last commit
Aug 2026
Advanced
Item type
skill
Key
masking-pii-data
Source
github.com/unknown-333/awesome-data-engineering-skills