multimodal-medical-imaging
SkillDev toolsThe Multimodal Medical Imaging Analysis Skill leverages state-of-the-art Vision-Language Models (VLMs) like Gemini 1.5 Pro and GPT-4o to interpret medical imagery alongside clinical text.
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the multimodal-medical-imaging skill
About this skill
The largest open-source medical AI skills library for OpenClaw🦞.
What this skill tells your AI
The instructions your AI receives, as published by freedomintelligence/openclaw-medical-skills in skills/multimodal-medical-imaging/SKILL.md and read by ahel’s review.
name: 'multimodal-medical-imaging' description: 'Analyzes medical images (X-ray, MRI, CT) using multimodal LLMs to identify anomalies and generate reports.' measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools:
- read_file
- run_shell_command
Multimodal Medical Imaging Analysis
The Multimodal Medical Imaging Analysis Skill leverages state-of-the-art Vision-Language Models (VLMs) like Gemini 1.5 Pro and GPT-4o to interpret medical imagery alongside clinical text.
When to Use This Skill
- When you need a preliminary screening of medical images.
- When correlating visual findings with textual clinical notes.
- To generate structured reports (DICOM-SR-like) from raw images.
Core Capabilities
- Anomaly Detection: Identify potential pathologies in X-rays, CTs, etc.
- Report Generation: Draft radiology reports in standard formats.
- VQA (Visual Question Answering): Answer specific questions about an image (e.g., "Is there a fracture in the left femur?").
Workflow
- Input: Provide an image file path (JPG, PNG) and a specific clinical question or "generate report" instruction.
- Analyze: The agent sends the image and prompt to the VLM.
- Output: Returns a JSON object with findings, confidence scores, and reasoning.
Example Usage
User: "Analyze this chest X-ray for pneumonia."
Agent Action:
python3 Skills/Clinical/Medical_Imaging/Multimodal_Analysis/multimodal_agent.py \
--image "/path/to/cxr.jpg" \
--prompt "Check for signs of pneumonia and consolidation."
Advanced
- Item type
- skill
- Key
multimodal-medical-imaging- Source
- github.com/freedomintelligence/openclaw-medical-skills
github.com/freedomintelligence/openclaw-medical-skills
Related picks
Skill · probabl-ai
The pick for Pythoncc-python-dev
Skill · doccker
The pick for Pythonremotion-markup
Skill · remotion-dev
More in Dev toolsremotion-interactivity
Skill · remotion-dev
More in Dev toolsimprove-codebase-architecture
Skill · mattpocock
More in Dev toolsgrilling
Skill · mattpocock
More in Dev tools