OpenEvidence Synthetic Workflow Evaluation Loop
SkillDev toolsEvaluate OpenEvidence product workflows safely with synthetic cases before any clinical rollout. Use when working with OpenEvidence in a healthcare organization. Trigger with "openevidence local dev loop", "OpenEvidence evaluation", or a matching workflow request.
Use OpenEvidence Synthetic Workflow Evaluation Loop in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add OpenEvidence Synthetic Workflow Evaluation Loop and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the OpenEvidence Synthetic Workflow Evaluation Loop skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; Ahel provides instructions and does not run this skill.
No other account needed.
Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
What this skill tells your AI
The instructions your AI receives, as published by jeremylongshore/tons-of-skills-marketplace in skills/.curated/openevidence-local-dev-loop/SKILL.md and read by Ahel’s review.
Overview
Replace the nonexistent local developer loop with a repeatable browser/app evaluation process. Keep inputs minimal, separate observed facts from assumptions, and leave consequential decisions with the named accountable owner.
Prerequisites
- A clearly bounded workflow, accountable clinical owner, and organizational policy
- Current first-party OpenEvidence documentation and applicable institution agreements
- Synthetic or properly authorized minimum-necessary data
Tool Discipline
Use Read, Glob, and Grep to inspect supplied policies, plans, and evidence. Use WebFetch only for current first-party OpenEvidence documentation. Use Write or Edit only when the user requests a named deliverable with an approved destination. Never expose credentials, PHI, recordings, or unrestricted environment output.
Current Contract
- OpenEvidence does not publish a local runtime, sandbox SDK, or public test API in the audited documentation.
- Synthetic scenarios are the default evaluation input; real PHI requires the full approved data boundary.
- Evaluation tests workflow fitness and evidence review, not medical-device validation.
Authentication
Use only the official OpenEvidence web/mobile sign-in or an institution-approved access path. Do not invent API keys, OAuth clients, SDK credentials, service accounts, or private endpoints. Never ask a user to reveal a password, session token, cookie, or recovery code.
Instructions
- Define feature, user role, expected workflow outcome, unacceptable failure, and clinical reviewer.
- Create synthetic scenarios spanning routine, ambiguous, conflicting-evidence, and failure-path cases.
- Read the current guide and record the model, feature, and surface used for each run.
- Execute manually through the supported product, preserving only de-identified prompts, citations, and observations.
- Have a qualified reviewer score traceability, applicability, uncertainty, and workflow burden.
- Iterate one variable at a time and publish a go, revise, or stop recommendation.
Approval Boundaries
Do not create or share accounts; change access, roles, agreements, consent, retention, or security settings; enter PHI; record a conversation; copy content into another system; contact a patient; make a diagnosis or treatment decision; submit billing; transmit a support packet; run a production pilot; or represent vendor capabilities without explicit approval from the accountable owner. A qualified professional remains responsible for clinical decisions.
Output
Return scope, current first-party evidence and date, data classification, workflow or findings, citations reviewed, assumptions rejected, clinical and governance owners, approval state, unresolved risk, and the exact next action. Redact patient and credential data.
Error Handling
| Condition | Response |
|---|---|
| No sandbox | Use synthetic inputs in an authorized account; do not probe private infrastructure. |
| Output non-deterministic | Score invariant qualities rather than exact wording. |
| Reviewer disagreement | Preserve both rationales and escalate to the clinical owner. |
Examples
This compact example shows the minimum reviewable handoff; adapt fields to the approved workflow without adding sensitive data.
Input:
feature=Ask; cases=12 synthetic; reviewer=clinical lead; surface=web
Expected handoff:
runs=12; acceptable=9; revise=2; stop=1; next-change=prompt
Resources
Signals
- GitHub stars
- 3k
- Forks
- 415
- Last commit
- Oct 2026
Advanced
- Item type
- skill
- Key
openevidence-local-dev-loop- Source
- github.com/jeremylongshore/tons-of-skills-marketplace
github.com/jeremylongshore/tons-of-skills-marketplace