Product Positioning Evaluation Framework
SkillDev toolsProduct Positioning Evaluation Framework
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Product Positioning Evaluation Framework skill
What this skill tells your AI
The instructions your AI receives, as published by markmhendrickson/neotoma in .claude/skills/evaluate_positioning/SKILL.md and read by ahel’s review.
This skill evaluates whether a positioning surface (homepage, subpage, pitch, README, or other asset) effectively communicates to a target ICP. It combines strategic positioning assessment with surface expression evaluation — because positioning only exists as experienced by the audience.
Five complementary lenses:
- Obviously Awesome (April Dunford) — The 5-step positioning process: competitive alternatives, unique attributes, value mapping, best-fit customers, market category.
- Made to Stick (Chip & Dan Heath) — The SUCCESs framework for evaluating whether positioning is memorable and actionable: Simple, Unexpected, Concrete, Credible, Emotional, Stories.
- Neotoma Evaluator Feedback — Ground-truth validation against all evaluator feedback stored in Neotoma, surfacing whether positioning claims match what real evaluators said, felt, and objected to.
- ICP Alignment — Structural fit against
docs/icp/primary_icp.md: pain triggers, operational modes, vocabulary bridge, qualification criteria, adoption funnel, and competitor migration paths. - Surface Expression — Whether the page or asset presenting the positioning supports comprehension, demonstrates the product, and follows the audience's natural decision flow.
Lenses 1-4 produce the weighted Combined Positioning Score. Lens 5 (Surface Expression) functions as a gate: if it scores below 6, the combined score carries an asterisk.
Core Principle
Positioning is not messaging. Positioning is context.
Positioning defines the context within which customers evaluate your product. It determines what category customers place you in, what alternatives they compare you against, which features they pay attention to, and how they judge your value. Get positioning right, and everything downstream — messaging, sales pitches, marketing campaigns, pricing — becomes dramatically easier. Get it wrong, and no amount of clever copywriting or advertising spend will save you. Customers who don't understand what you are will never understand why you matter.
The foundation of great positioning is understanding that customers always evaluate products relative to alternatives. There is no such thing as absolute product perception. A product that seems expensive in one context seems cheap in another. A feature that seems innovative against one set of competitors seems table-stakes against another. Your job is to deliberately choose the context that makes your unique strengths obvious.
Scoring Overview
Goal: 10/10 — Rate positioning quality across four lenses. The Obviously Awesome (Dunford) score below is one of four components; see Combined Positioning Score for the weighted composite.
Obviously Awesome Score (Dunford): 0-10
| Score | Description |
|---|---|
| 0-2 | No clear positioning. Customers can't explain what the product is or who it's for. |
| 3-4 | Vague positioning. Category is unclear, differentiation is weak, target customer is "everyone." |
| 5-6 | Partial positioning. Some components are clear but others are missing or inconsistent. Team members describe the product differently. |
| 7-8 | Strong positioning. All five components are defined. Team is aligned. Customers generally understand the value. |
| 9-10 | Exceptional positioning. Every component reinforces the others. Customers immediately understand the product, why it's different, and why they should care. The positioning creates an "aha" moment. |
Additional scores: Made to Stick (0-10), Evaluator Feedback Grounding (0-10), ICP Alignment (0-10). See each section below for criteria.
Made to Stick Evaluation (SUCCESs Framework)
The SUCCESs framework from Chip & Dan Heath's "Made to Stick" evaluates whether an idea — including product positioning — is structured to be understood, remembered, and acted upon. Positioning that scores well on Dunford's framework but fails the stickiness test won't survive first contact with the market.
The 6 SUCCESs Principles
Each principle is scored 0-10. The Made to Stick composite is the average.
1. Simple — Find the Core
What it means: Strip the positioning to its essential core. Not dumbing down — finding the single most important thing customers need to understand. If you can't say it in one sentence, it's not simple enough.
Evaluation criteria:
- Can the positioning be expressed as a single, compact statement?
- Does it use the "inverted pyramid" — lead with the most important insight?
- Is there a "Commander's Intent" — a guiding principle that survives contact with the real world?
- Does it avoid burying the lead under technical architecture or feature lists?
- Core Claim Coherence: Is the Commander's Intent reinforced across all positioning elements (hero, value props, proof points, competitive framing, CTA), not just stated once? A core claim that appears only in the headline but isn't echoed in the rest of the surface loses its force.
Scoring:
| Score | Evidence |
|---|---|
| 0-3 | Positioning requires multiple paragraphs or a diagram to explain. Technical jargon dominates. Core claim (if present) appears in one place only. |
| 4-6 | Core idea exists but is buried under qualifications, features, or abstractions. Some reinforcement across surface but inconsistent. |
| 7-8 | Clear one-sentence core. Most people would remember it after hearing it once. Core claim echoed in at least 3 surface elements. |
| 9-10 | Commander's Intent is obvious. Anyone on the team could say it the same way. It's a proverb-level compression of the value. Every surface section reinforces the same core — the page has one voice. |
Common failure mode for technical products: Leading with architecture ("append-only observation log with schema-bound entities") instead of felt experience ("your agent never forgets, and you can prove it").
2. Unexpected — Get and Hold Attention
What it means: Break a pattern to get attention, then create a curiosity gap to hold it. Positioning should violate the audience's expectations in a way that makes them want to know more.
Evaluation criteria:
- Does the positioning break the audience's existing schema or expectations?
- Does it open a curiosity gap — a question the audience needs answered?
- Does it avoid the "so what?" reaction from the target customer?
- Is the unexpected element connected to the core value (not just shock for shock's sake)?
Scoring:
| Score | Evidence |
|---|---|
| 0-3 | Positioning is predictable. Reads like every other tool in the category. No reason to stop scrolling. |
| 4-6 | Has a twist or insight, but it doesn't generate a curiosity gap. Audience nods but doesn't lean in. |
| 7-8 | Clear pattern violation that makes the audience pause. Creates a "wait, what?" moment connected to real pain. |
| 9-10 | Reframes the problem space in a way the audience hadn't considered. They can't stop thinking about the gap it opened. |
Diagnostic question: "After hearing the positioning for 5 seconds, does the target customer want to hear more, or do they think they already know what this is?"
3. Concrete — Make It Real
What it means: Use sensory, specific language instead of abstractions. People remember concrete images and specific examples, not abstract principles.
Evaluation criteria:
- Does the positioning use specific, tangible examples rather than abstract concepts?
- Can the audience picture a specific scenario or use case?
- Does it avoid vague benefits ("better productivity," "more efficient") in favor of specific outcomes ("your agent recalls the correction you made last Tuesday")?
- Does it use the audience's own vocabulary (see ICP Vocabulary Bridge)?
- Mental Model Anchoring: For new-category products, does the positioning anchor the unfamiliar concept to a system the audience already understands? The analogy must map accurately without creating false expectations. (Example: "Git for agent state" maps versioning → commits, diff → change inspection, replay → checkout.)
Scoring:
| Score | Evidence |
|---|---|
| 0-3 | All abstract. "Deterministic state layer" with no grounding. Audience can't picture using it. No familiar reference point for new concepts. |
| 4-6 | Mix of abstract and concrete. Some examples exist but they feel generic or hypothetical. Analogy attempted but mapping is loose or misleading. |
| 7-8 | Specific scenarios the audience recognizes from their own experience. Uses their vocabulary. For new categories: effective anchoring analogy that maps accurately. |
| 9-10 | Velcro-level concreteness — multiple hooks that stick. The audience can immediately picture the before/after in their own workflow. Anchoring analogy is so apt the audience uses it to explain the product to others. |
The Velcro Theory of Memory: The more sensory hooks an idea has, the more it sticks. Abstract positioning has smooth surfaces; concrete positioning is covered in hooks.
4. Credible — Enable Belief
What it means: Make the positioning believable without requiring trust. External credibility (authorities, data) helps, but the most powerful form is testable credibility — letting the audience verify the claim themselves.
Evaluation criteria:
- Does the positioning include verifiable proof points (customer quotes, data, case studies)?
- Can the audience test the claim themselves without a large investment (the "try before you trust" test)?
- Does it use internal credibility — vivid details that signal genuine experience?
- Does it avoid claims that require the audience to take the company's word for it?
Scoring:
| Score | Evidence |
|---|---|
| 0-3 | Claims with no evidence. "We're the best" with nothing to verify. |
| 4-6 | Some proof points but they're generic or unverifiable. Testimonials feel curated. |
| 7-8 | Specific, verifiable claims backed by data or testable within a short evaluation. |
| 9-10 | "Sinatra Test" passed — if it works for [impressive case], it'll work for me. Plus self-testable via eval or install. |
The Sinatra Test: "If I can make it there, I can make it anywhere." One compelling case study that makes the audience extrapolate credibility.
5. Emotional — Make People Care
What it means: Appeal to self-interest and identity, not just logic. People act when they feel something. For technical products, the emotional lever is often identity-based: "people like me use this" or "this is the kind of tool a serious builder uses."
Evaluation criteria:
- Does the positioning connect to the audience's identity (who they want to be)?
- Does it tap into self-interest at the right level (not just "save money" but "become the builder who doesn't babysit agents")?
- Does it use the "role transformation" framing (escaping → into)?
- Does it make the audience feel understood rather than sold to?
Scoring:
| Score | Evidence |
|---|---|
| 0-3 | Pure logic/features. No emotional resonance. Could be a spec sheet. |
| 4-6 | Some identity appeal but it feels performative or misaligned with the audience's self-image. |
| 7-8 | Clear identity hook. Audience thinks "this was built for people like me." Role transformation is visible. |
| 9-10 | The positioning feels personal. It names a pain the audience has felt but hasn't articulated. They share it because it says something about them. |
Technical product trap: Over-indexing on logic because "our audience is rational." Even engineers choose tools based on identity and community — they just don't admit it.
6. Stories — Drive Action Through Narrative
What it means: Stories simulate experience and inspire action. The best positioning embeds a story — even implicitly — that helps the audience see themselves using the product and achieving the outcome.
Evaluation criteria:
- Does the positioning contain or imply a narrative arc (before → transformation → after)?
- Can the audience project themselves into the story?
- Does it use the challenge plot (overcoming obstacles), connection plot (bridging gaps), or creativity plot (novel solution)?
- Does it show, not tell?
Scoring:
| Score | Evidence |
|---|---|
| 0-3 | No narrative. Just claims and features. |
| 4-6 | Implicit story exists but it's buried. The before/after isn't vivid. |
| 7-8 | Clear narrative the audience recognizes. They can see the transformation. Sales narrative (problem → old way → new way → solution → proof) is embedded. |
| 9-10 | The positioning IS a story. Every element contributes to a narrative the audience wants to be part of. They retell it to others. |
Made to Stick Composite Score
Average the six principle scores. Map to positioning quality:
| Composite | Interpretation |
|---|---|
| 0-3 | Positioning won't survive first contact. Forgettable. |
| 4-5 | Some sticky elements but the overall message doesn't cohere. |
| 6-7 | Positioning sticks with the right audience but has weak spots that dilute impact. |
| 8-9 | Strong stickiness across most dimensions. Message travels well. |
| 10 | Proverb-level. The positioning becomes the way people talk about the category. |
Neotoma Evaluator Feedback Grounding
Purpose: Ground positioning evaluation in real-world feedback data from Neotoma evaluators. This section turns positioning assessment from a theoretical exercise into an evidence-based audit.
Retrieval Protocol
When running a positioning evaluation, retrieve evaluator feedback from Neotoma before scoring:
- Retrieve evaluators: Query
developer_release_testerentities to get the full evaluator roster. - Retrieve feedback: Query
issueentities (and linked conversations) for evaluator reactions, objections, and quotes. - Retrieve evaluator-to-ICP mappings: Search for
agent_messageentities containing evaluator ICP scoring and classification (Strong, Moderate, Weak, Non-ICP, DIY). - Retrieve convergent themes: Search for agent messages about evaluation analysis, positioning gaps, and feedback synthesis.
Evaluation Dimensions
Score each dimension 0-10 based on evaluator evidence:
A. Pain Validation (0-10)
Does the positioning name a pain that evaluators actually described? Does it lead with that pain?
| Score | Evidence required |
|---|---|
| 0-3 | Positioning describes a pain no evaluator mentioned. Theoretical problem. Capabilities-first framing. |
| 4-6 | Pain exists in evaluator data but positioning frames it differently than evaluators describe it. Pain is present but buried below features or architecture. |
| 7-8 | Multiple evaluators independently described the exact pain the positioning names. Surface leads with recognizable problems before introducing capabilities. |
| 9-10 | Evaluators used the same language the positioning uses. Verbatim alignment between positioning claims and evaluator quotes. Pain-first ordering: the audience recognizes the problem before they learn what the product does. |
Pain Ordering (product-type modifier): The form of "leading with pain" varies by product type. Load docs/icp/primary_icp.md to determine which framing applies:
- Infrastructure products: Lead with production failure modes (state drift, silent mutation, irreproducible decisions). Value framed as guarantees under failure conditions ("what cannot go wrong"), not capabilities.
- SaaS/workflow products: Lead with workflow pain (wasted time, lost context, manual workarounds). Value framed as outcomes ("what you can now do").
- Platforms: Lead with ecosystem friction (integration failures, lock-in, capability gaps). Value framed as ecosystem leverage ("what this unlocks").
Evaluator evidence themes to check against:
- Re-prompting / context-janitor tax (universal across evaluators)
- State drift between sessions and tools
- "What counts as a fact worth remembering?" cold-start confusion
- Build-in-house state workarounds (flat files, JSON, markdown, custom MCP+Postgres)
- Cross-tool state fragmentation
- JSON/file scaling limits triggering compensatory tooling
- "State integrity, not retrieval quality" as the differentiated layer
B. Competitive Frame Accuracy (0-10)
Does the positioning compare against what evaluators actually use today, or against strawmen?
| Score | Evidence required |
|---|---|
| 0-3 | Positioning compares against tools evaluators don't use (VC-funded competitors no one tried). |
| 4-6 | Some competitive alternatives match evaluator reality but others are hypothetical. |
| 7-8 | Competitive frame matches evaluator-reported alternatives: SQLite, git+markdown, flat JSON, Notion, platform memory, custom Postgres. |
| 9-10 | Positioning anticipates evaluator migration paths (the specific thing they tried, why it failed, what they're receptive to). |
Real competitive alternatives from evaluator data:
- Markdown/flat files (CLAUDE.md, SOUL.md, HEARTBEAT.md)
- JSON/CSV files that trigger compensatory Python scripts
- Notion/Airtable (breaks when adding agent reviewers)
- Platform memory (Claude, ChatGPT — tool-specific, non-auditable)
- Custom MCP + Postgres (capable DIY builders)
- SQLite (valid starting point but lacks versioning, conflict detection, provenance)
- RAG/vector memory (Mem0, Zep, LangChain — re-derives structure every session)
- "Do nothing" / raw re-prompting
C. Objection Coverage (0-10)
Does the positioning preemptively address objections evaluators actually raised?
| Score | Evidence required |
|---|---|
| 0-3 | Positioning ignores or is blind to common evaluator objections. |
| 4-6 | Some objections addressed but major ones are missing. |
| 7-8 | Core objections covered. Evaluators would feel heard. |
| 9-10 | Every major objection class has a clear, satisfying response embedded in or derivable from the positioning. |
Evaluator-sourced objection classes:
- "How is this different from RAG memory?" (most common confusion)
- "How is this better than SQLite?" (from technically sophisticated evaluators)
- "Should I use this alongside platform memory?" (coexistence question)
- "What counts as a fact worth remembering?" (cold-start/onboarding)
- "This feels like a solution looking for a problem" (from non-ICP or weak-fit evaluators)
- Supply chain / dependency security concerns (trust barrier)
- "Not my biggest problem right now" (from capable DIY builders)
- Architecture-first language that doesn't lead with felt experience
D. Evaluator Segment Coverage (0-10)
Does the positioning resonate with the evaluator segments that matter most?
| Score | Evidence required |
|---|---|
| 0-3 | Positioning resonates only with non-ICP evaluators or no segment. |
| 4-6 | Resonates with one segment but alienates or confuses others. |
| 7-8 | Strong resonance with "Strong ICP" evaluators. Moderate evaluators can see themselves in it. |
| 9-10 | Strong ICP evaluators would share it. Moderate evaluators would try it. Even Capable DIY evaluators recognize the problem description. |
Evaluator distribution for reference (as scored against primary_icp.md Q1-Q11/D1-D12):
- Strong ICP (~8): Multi-agent stack operators, autonomous pipeline builders, personal OS constructors
- Capable DIY (~2): Validate the problem strongly but building their own infrastructure
- Moderate (~7): Partial fit — some qualification criteria met, potential to convert with better activation
- Weak (~4): Low qualifier scores, not experiencing the core pain
- Non-ICP (~2): Hard disqualifiers (human-driven thought-partner pattern, state management is their product)
Feedback Grounding Composite Score
Average the four dimension scores (A-D). This score reflects how well positioning is grounded in reality versus theory.
ICP Alignment Evaluation
Purpose: Score positioning against the structural elements of docs/icp/primary_icp.md to ensure every positioning component serves the defined target customer.
Retrieval Protocol
Load docs/icp/primary_icp.md before scoring. Cross-reference with docs/icp/profiles.md for detailed profile data and docs/icp/developer_release_targeting.md for release-specific targeting constraints.
Evaluation Dimensions
Score each dimension 0-10:
I. Archetype Resonance (0-10)
Does the positioning speak to the primary ICP archetype: personal agentic OS builders/operators who spend significant effort compensating for the absence of reliable agent state?
| Score | Evidence |
|---|---|
| 0-3 | Positioning addresses a generic developer audience. No specificity to the agentic OS builder. |
| 4-6 | Partially addresses the archetype but misses key characteristics (multi-agent stacks, cross-session persistence, personal OS construction). |
| 7-8 | Clearly addresses someone who wires together multi-agent stacks and experiences state drift. |
| 9-10 | The ICP reads the positioning and thinks "this was written by someone who lives my workflow." The archetype section of primary_icp.md could be quoted as evidence. |
II. Pain Trigger Alignment (0-10)
Does the positioning connect chronic tax (re-prompting, manual sync) to acute crisis (corrupted state, bad decisions from wrong data)?
| Score | Evidence |
|---|---|
| 0-3 | Positioning mentions neither chronic tax nor acute crisis. Leads with features or architecture. |
| 4-6 | Addresses one pain trigger but not both, or doesn't connect them. |
| 7-8 | Both chronic and acute triggers present. The connection is clear: convenience opens the door, integrity closes the sale. |
| 9-10 | Uses the exact framing: "You're already paying the re-prompting tax every day. But the real risk isn't the time you waste — it's the time you don't notice your agent is operating on bad state." |
III. Operational Mode Coverage (0-10)
Does the positioning resonate across all three operational modes (Operating, Building, Infrastructure debugging)?
| Score | Evidence |
|---|---|
| 0-3 | Positioning addresses only one mode or none specifically. |
| 4-6 | Primary mode is addressed but the other modes feel like afterthoughts. |
| 7-8 | All three modes are served. The role transformation (escaping → into) is visible for each. |
| 9-10 | Each mode has a clear "that's me" moment. The positioning canvas works regardless of which mode the reader is in right now. |
Role transformations to check:
- Operating: context janitor → operator with continuity
- Infrastructure debugging: log archaeologist → platform engineer with replayable state
- Building: inference babysitter → builder on solid ground
IV. Vocabulary Bridge Compliance (0-10)
Does the positioning use the ICP's language, not internal/technical jargon?
| Score | Evidence |
|---|---|
| 0-3 | Internal jargon dominates: "deterministic state layer," "append-only observation log," "schema-bound entities." |
| 4-6 | Mix of ICP and internal language. Some bridging exists but technical terms leak through. |
| 7-8 | Leads with ICP vocabulary ("memory," "forgetting," "context," "what it knows"). Introduces technical terms only after bridging. |
| 9-10 | The ICP Vocabulary Bridge table from primary_icp.md is fully respected. Positioning could be understood by someone who says "fact" instead of "entity" and "memory" instead of "state." |
V. Qualification & Disqualification Clarity (0-10)
Does the positioning naturally attract qualified prospects and repel non-ICP?
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 32
- Forks
- 3
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
evaluate-positioning- Source
- github.com/markmhendrickson/neotoma