OpenEvidence Synthetic Workflow Evaluation Loop

SkillDev tools

Evaluate OpenEvidence product workflows safely with synthetic cases before any clinical rollout. Use when working with OpenEvidence in a healthcare organization. Trigger with "openevidence local dev loop", "OpenEvidence evaluation", or a matching workflow request.

Use OpenEvidence Synthetic Workflow Evaluation Loop in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add OpenEvidence Synthetic Workflow Evaluation Loop and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the OpenEvidence Synthetic Workflow Evaluation Loop skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

OpenEvidence Synthetic Workflow Evaluation LoopStart free

What this skill tells your AI

The instructions your AI receives, as published by jeremylongshore/tons-of-skills-marketplace in skills/.curated/openevidence-local-dev-loop/SKILL.md and read by Ahel’s review.

Overview

Replace the nonexistent local developer loop with a repeatable browser/app evaluation process. Keep inputs minimal, separate observed facts from assumptions, and leave consequential decisions with the named accountable owner.

Prerequisites

  • A clearly bounded workflow, accountable clinical owner, and organizational policy
  • Current first-party OpenEvidence documentation and applicable institution agreements
  • Synthetic or properly authorized minimum-necessary data

Tool Discipline

Use Read, Glob, and Grep to inspect supplied policies, plans, and evidence. Use WebFetch only for current first-party OpenEvidence documentation. Use Write or Edit only when the user requests a named deliverable with an approved destination. Never expose credentials, PHI, recordings, or unrestricted environment output.

Current Contract

  • OpenEvidence does not publish a local runtime, sandbox SDK, or public test API in the audited documentation.
  • Synthetic scenarios are the default evaluation input; real PHI requires the full approved data boundary.
  • Evaluation tests workflow fitness and evidence review, not medical-device validation.

Authentication

Use only the official OpenEvidence web/mobile sign-in or an institution-approved access path. Do not invent API keys, OAuth clients, SDK credentials, service accounts, or private endpoints. Never ask a user to reveal a password, session token, cookie, or recovery code.

Instructions

  1. Define feature, user role, expected workflow outcome, unacceptable failure, and clinical reviewer.
  2. Create synthetic scenarios spanning routine, ambiguous, conflicting-evidence, and failure-path cases.
  3. Read the current guide and record the model, feature, and surface used for each run.
  4. Execute manually through the supported product, preserving only de-identified prompts, citations, and observations.
  5. Have a qualified reviewer score traceability, applicability, uncertainty, and workflow burden.
  6. Iterate one variable at a time and publish a go, revise, or stop recommendation.

Approval Boundaries

Do not create or share accounts; change access, roles, agreements, consent, retention, or security settings; enter PHI; record a conversation; copy content into another system; contact a patient; make a diagnosis or treatment decision; submit billing; transmit a support packet; run a production pilot; or represent vendor capabilities without explicit approval from the accountable owner. A qualified professional remains responsible for clinical decisions.

Output

Return scope, current first-party evidence and date, data classification, workflow or findings, citations reviewed, assumptions rejected, clinical and governance owners, approval state, unresolved risk, and the exact next action. Redact patient and credential data.

Error Handling

ConditionResponse
No sandboxUse synthetic inputs in an authorized account; do not probe private infrastructure.
Output non-deterministicScore invariant qualities rather than exact wording.
Reviewer disagreementPreserve both rationales and escalate to the clinical owner.

Examples

This compact example shows the minimum reviewable handoff; adapt fields to the approved workflow without adding sensitive data.

Input:

feature=Ask; cases=12 synthetic; reviewer=clinical lead; surface=web

Expected handoff:

runs=12; acceptable=9; revise=2; stop=1; next-change=prompt

Resources

Signals

GitHub stars
3k
Forks
415
Last commit
Oct 2026
Advanced
Item type
skill
Key
openevidence-local-dev-loop
Source
github.com/jeremylongshore/tons-of-skills-marketplace