call-cognitive-load-monitor

SkillMonitoring & ops

Offline experimental CALL-E transcript heuristics for repetition, confusion phrases and conversational proxies, with advisory scores and script suggestions for human review; not a cognitive assessment or consent determination.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the call-cognitive-load-monitor skill

What this skill tells your AI

The instructions your AI receives, as published by calle-ai/awesome-phone-call-agents in skills/call-cognitive-load-monitor/SKILL.md and read by ahel’s review.

Detect when your caller is overwhelmed — before it costs you the call.

This skill analyzes supplied text for heuristic confusion markers and conversational proxies. It does not measure acoustics, administer a cognitive assessment, or establish consent validity. Reports and suggested script changes require human review; no live integration, message sending or follow-up call is performed.


Why This Skill Exists

Most call-centre quality tools measure whether the agent performed well. This skill measures whether the caller was able to process what was said. When cognitive load peaks — especially at closing, where consent and commitment are given — the validity of that interaction is at risk.

Advisory scope: A CONSENT_AT_RISK label is a suggestion for human review, not a regulatory finding or authorization for a follow-up call. No label establishes that consent is valid or invalid.


Research Scope

The weights and thresholds are unvalidated demonstration choices, not a reproduction of a published acoustic model. Background and claim limitations: references/research-papers.md.


Quick Start

python3 scripts/monitor_cognitive_load.py \
  --transcript path/to/transcript.json \
  --dry-run \
  --out /tmp/load_report.json

Validate output schema

python3 scripts/validate_load_report.py --report /tmp/load_report.json

Input

Standard CALL-E transcript — same two formats as other skills:

Format A — array of turns:

[
  {"role": "agent",  "text": "Under the terms of sub-clause 4(b)(ii) of the agreement..."},
  {"role": "callee", "text": "Sorry, could you say that again? I'm not following."},
  {"role": "agent",  "text": "Of course. Regarding the billing section specifically..."},
  {"role": "callee", "text": "I still don't understand. What does that mean for me?"}
]

Output — Overload Report

{
  "call_id": "calle-001",
  "analysis_timestamp": "2026-09-17T08:00:00Z",
  "overall_cognitive_load": "HIGH",
  "load_score": 0.78,
  "load_by_phase": {
    "opening":  0.12,
    "middle":   0.45,
    "closing":  0.78
  },
  "peak_phase": "closing",
  "overload_signals": [
    {
      "type": "repetition_request",
      "evidence": "Sorry, could you say that again?",
      "turn": 4,
      "phase": "middle",
      "weight": 0.30
    },
    {
      "type": "confusion_phrase",
      "evidence": "I still don't understand. What does that mean for me?",
      "turn": 6,
      "phase": "middle",
      "weight": 0.25
    },
    {
      "type": "jargon_density_spike",
      "evidence": "sub-clause 4(b)(ii) of the agreement",
      "turn": 3,
      "phase": "middle",
      "weight": 0.18
    }
  ],
  "interaction_dynamics": {
    "turn_imbalance_score": 0.62,
    "avg_silence_gap_ms": 2800,
    "overlap_count": 0,
    "participation_ratio": {"agent": 0.74, "callee": 0.26}
  },
  "script_patches": [
    {
      "original": "Under the terms of sub-clause 4(b)(ii) of the agreement",
      "suggested": "According to our billing rules",
      "rationale": "Legal jargon removed; Flesch-Kincaid grade reduced from 18 to 6"
    }
  ],
  "consent_validity_flag": "AT_RISK",
  "consent_validity_reason": "Load score 0.78 exceeded threshold 0.65 during closing phase where commitment was recorded.",
  "recommended_action": "SEND_WRITTEN_CONFIRMATION",
  "false_positive_disclaimer": "Cognitive load inference is probabilistic. A flagged call does not constitute a legal finding. Human review is required before any adverse action.",
  "flags": ["REQUIRES_HUMAN_REVIEW", "JARGON_DENSITY_HIGH"],
  "analysis_mode": "heuristic",
  "schema_version": "1.0"
}

Cognitive Load Signals

Linguistic Markers

Signal TypeExampleWeight
repetition_request"Could you repeat that?" / "Say that again?"0.30
confusion_phrase"I don't understand" / "I'm lost" / "What does that mean?"0.25
self_correction"Wait, I mean... actually..."0.15
clarification_request"So what you're saying is...?"0.20
jargon_density_spikeTechnical / legal terminology per 50-word window0.18
hedge_word_cluster"I think", "maybe", "I'm not sure" (3+ in one turn)0.12

Interaction Dynamic Markers

Signal TypeThresholdSignificance
turn_imbalanceAgent:Callee ratio > 3:1Agent dominating — callee cannot process
long_silence_gap> 2 500 msCallee processing delay — indicates high extraneous load
participation_dropCallee turns drop ≥ 40% in closing phaseCognitive withdrawal

Load Levels

Load ScoreLevelRecommended Action
< 0.35LOWPROCEED
0.35–0.64MEDIUMFLAG_FOR_REVIEW
0.65–0.84HIGHSEND_WRITTEN_CONFIRMATION
≥ 0.85CRITICALREPEAT_CALL_WITH_SIMPLER_SCRIPT

Consent Validity Flags

FlagTriggerMeaning
CONSENT_VALIDLoad < 0.65 during closingCaller was processing normally when consenting
CONSENT_AT_RISKLoad ≥ 0.65 during closingCaller may have consented under cognitive strain
CONSENT_UNKNOWNClosing phase not identifiableTranscript structure too short to assess

[!CAUTION] CONSENT_AT_RISK does not mean consent is invalid. It flags the call for human review and recommends sending written confirmation. Never use this flag as a standalone legal determination.


Command-Line Reference

usage: monitor_cognitive_load.py [-h] --transcript TRANSCRIPT
                                 [--threshold THRESHOLD]
                                 [--dry-run]
                                 [--out OUT]

options:
  --transcript   Path to the transcript JSON file (required)
  --threshold    Load score above which level is HIGH (default: 0.65)
  --dry-run      Analyse without side effects
  --out          Write report JSON to this path (default: stdout)

Integration with CALL-E

Insert as a post-call QA step in any pipeline where consent or commitment is recorded:

[call ends] → [transcribe] → [monitor_cognitive_load.py]
                                        ↓
                   load=HIGH → send written confirmation email
                   load=CRITICAL → schedule follow-up call with simplified script

Combine with client-persona-profiler: Analytical (C) and Steady (S) callers have lower jargon tolerance — flag their calls at a lower threshold.


Privacy & Safety

  • Reports contain copied transcript evidence and script-patch text and may contain private data. Keep real-input reports local and access-controlled; use synthetic inputs for examples and tests. Do not publish reports without review.
  • The separate validate_load_report.py check is not automatically run before writing and is not a complete PII redactor or privacy guarantee.
  • The consent-validity flag is advisory only. A human must review before any regulatory or legal action.

Full safety reference: references/safety.md

Expected Outcomes & Metrics

The following numbers are unvalidated design targets, not measured results.

MetricExpected TargetNotes
Latency< 50ms per transcriptEvaluates locally without LLM dependencies.
TPR (True Positive Rate)> 85%Identifies cognitive strain in human-annotated datasets.
FPR (False Positive Rate)< 10%Some benign clarification requests may be flagged.

Limitations & Known Constraints

  • Text-Only Modality: Cannot detect sighs, actual silence duration, or exasperated tone. Silence and load fields are text-based proxies, not acoustic measurements.
  • Prototype options: --threshold is currently accepted but unused; reported levels use built-in constants. Short transcripts and a CONSENT_VALID label do not prove that a closing phase or consent occurred.
  • ASR Dependency: If the ASR mistranscribes "I'm lost" as "I boss", the signal is missed.
  • Language Bias: Heuristics are currently calibrated exclusively for English.

Files

skills/call-cognitive-load-monitor/
├── SKILL.md
├── scripts/
│   ├── monitor_cognitive_load.py      ← Main analysis runner
│   ├── validate_load_report.py        ← Output schema validator
│   └── test_cognitive_load.py         ← Test suite (50+ assertions)
└── references/
    ├── example-transcript.json        ← High-load example transcript
    ├── examples.md                    ← Usage examples
    ├── research-papers.md             ← Full citations
    └── safety.md                      ← Consent-validity ethics guide

Signals

GitHub stars
104
Forks
527
Last commit
Sep 2026
Advanced
Item type
skill
Key
call-cognitive-load-monitor
Source
github.com/calle-ai/awesome-phone-call-agents