Clone prototype
SkillFiles & storageClone a real app's screens as pixel-accurate, self-contained HTML artboards on the prototype canvas. Overlay a grid on the reference and sample colours visually, derive one measured design-token block, generate one HTML file per screen from a single script, verify by re-rendering, and park the reference underneath its mockup. Use when asked to 100% copy / clone an app's UI, rebuild screens from screenshots or Mobbin, extract a design system from reference images, or check a mockup against its reference.
Available today. Use it from your connected AI after setup.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
Then ask your AI: use the Clone prototype skill
What this skill tells your AI
The instructions your AI receives, as published by resciencelab/super-prototyping in skills/sp-clone-prototype/SKILL.md and read by ahel’s review.
Seven phases, in order. Never skip ahead. Tokens before HTML, sampling before tokens. Every colour and every metric in the final HTML must trace back to a measurement, not to a guess that "looks about right". The one thing built out of order is Phase 5's reference row: it needs no measurement, so it goes up first (see Phase 5).
Output lands in canvases/<slug>/ and the canvas picks it up automatically.
The sp-canvas skill covers how folders become tldraw pages and how
layout.json rows work. Name the folder for the source, e.g. notion-ios.
Toolkit: refkit (grid / sample / bands / bbox / scan / hairline / font /
shoot / diff / blend / tokens / batch / ink / crops / key / montage), plus
artgen for the rare asset that has to be drawn. Both are commands on your
PATH; shoot additionally needs Google Chrome (on Windows, Edge will do).
refkit --help
sp, refkit or artgen not found? Run sh <sp-canvas skill dir>/scripts/install.sh (Windows: install.ps1)
and follow that skill's references/install.md. A [super-prototyping:notice]
line on their stderr carries its own rule: finish the step, then do what it says.
Before Phase 1, check the project ignores what a run produces, once per
project: a run writes third-party captures into ref-*.html and
assets/refs/, and committing those is hard to undo. Work in
canvases/<slug>/scratch/. Each rule is checked on its own, and the
**/ matters — a pattern with a slash in the middle only matches at the root:
for p in 'ref-*.html' '**/assets/refs/' 'scratch/'; do
grep -qxF "$p" .gitignore 2>/dev/null || echo "$p" >> .gitignore
done
Then name the board folder once and address everything through it: captures
under $B/assets/refs/, everything a run derives under $B/scratch/. A -o
without a directory writes into whatever the current directory happens to be,
which is the user's project root as often as not:
B=canvases/<slug>
The folder skeleton ships with the app, which is installed outside your project. Address it through the kit root:
KIT="$(sp root)"
Finished runs to copy from are community projects, not bundled with the app: sp fetch <id> prints the folder of a read-only copy, its boards at
<that folder>/canvases/<slug>/. Phase 2 and Phase 6 below name the ids.
Phase 0: collect references
Get the highest-resolution capture you can; every later measurement is capped by it.
- The user's own screenshot is usually the authority on which screens
and which scroll state. Save each one as its own crop (
p1.png … pN.png) in$B/assets/refs/before doing anything else. Image caches rotate and the attachment will disappear mid-task. - Mobbin MCP (
mcp__mobbin__search_screens,search_flows), when available. Run one search per screen,platform: "ios", and keeptask_intentidentical across every call in the run. Describe the screen in plain language including its literal on-screen strings; that is what actually matches. Useexclude_screen_ids(a JSON array of quoted UUID strings) to push past near-misses you already rejected. Download withcurl -sL <image_url>. Results are ~299 × 678 webp, fine for placing on the canvas but too small to be the only sampling source. Cite results as markdown links to theirmobbin_url. - A native-resolution capture of any screen from the same app (@2x/@3x, e.g. 1179 × 2556) settles ink, scrim and accent values that a downscaled strip cannot. It does not have to be one of the screens you are cloning.
Check the colour space before you sample anything
A capture straight off a device is often untagged Display P3, and every tool in this kit reads raw bytes. Sampled as-is, a P3 capture and an sRGB one of the same screen disagree by 5-10 levels on any saturated colour, and nothing about either looks wrong on its own. Convert to sRGB first, and convert the whole set, so one token cannot end up averaging two spaces.
The test is cheap and it is the only thing that finds this: sample one
element that appears in every capture (a brand mark, an accent button, the
page ground) and compare across batches. Values that split into two clusters
are two colour spaces, not two colours. The claude-ios run had captures
01-07 in sRGB and 08-15 untagged P3; the same brand orange read #E07A54 on
one half and #D97757 on the other, and the page ground split with it. One
sips/Pillow conversion pass up front collapses both.
Two oranges may still survive the conversion, and then they are real: that run
kept #D97757 for the star mark and #CB6442 for the send button, on the same
screen. Convert first, then decide what is one token and what is two.
Record the capture scale once, in capture px per design pt, and reuse it everywhere:
scale = screen_inner_width_px / device_pt_width # e.g. 300 / 393 = 0.7634
Cross-check against height (649 / 852 = 0.7617). If the two disagree by
more than ~1%, the crop is wrong. Recrop before sampling.
Phase 1: grid on the image, then LOOK
Draw the grid onto the pixels and read the result with your eyes. That is the rule that makes this work. Sampling coordinates blind produces numbers with no idea which UI element they belong to, and those numbers end up in the wrong token.
refkit grid "$B/assets/refs/p4.png" \
-o "$B/scratch/g04.png" --zoom 3 --minor 10 --major 50
Then read g04.png as an image. Cyan every 10 source px, red and
labelled every 50. Walk the screen element by element and write down the
region each one occupies in source coordinates. Only then sample.
references/measuring.md is the technique: the
four kinds of colour region and the command each one needs, how gutters, row
heights and radii come off the same grid, and how to measure the type face
instead of guessing it. Read it before you sample anything.
Deliverable of this phase
A table with an evidence column, one row per token. Anything without evidence does not become a token.
| token | value | evidence |
|---|---|---|
--n-font | -apple-system, "SF Pro"… | refkit font on the page title, 0.93 (2nd 0.87) |
--n-text | #2C2C2C | H1 core ink @3x, 3 screens |
--n-hairline | #E9E8E7 | 1pt coverage solve, settings dividers |
--n-border | #EFEEEC | card outline solve |
Write the machine half of that table as you measure: probes.json in the
canvas folder, committed with the boards, one entry per measurement, {"id", "img", "cmd", "box"}.
Phase 4 replays it against your renders, so every number that justified a
token is re-checked after the token changes. Without it the evidence
drifts: one finished run shipped three rows still citing an alpha of .174
after the token had become .10, and refkit tokens cannot catch that; it
checks that tokens exist, not that their evidence still agrees with them.
Two things batch will not tell you. It compares the first colour a probe
prints, and sample prints its flat-fill census first, so an ink probe
needs --only ink or it silently compares two backgrounds and reports a
perfect Δ 0. And a key starting with _ in a probe is ignored, which is
where the sanity note below lives.
Every probe box carries a one-line sanity note proving the window holds
the element and only it. The window is wrong far more often than the
measurement is: a probe at cy=681 for a button row that sits at 564, or a
box that catches the neighbouring label instead of the glyph, returns a
confident, plausible number either way.
Re-sample rather than argue. Where the strip and a native capture disagree, the native capture wins.
Phase 2: design system before any screen
Write one :root block and inline it, byte-identically, into every
artboard. Artboards render in <iframe srcDoc sandbox="">, so there is no
shared stylesheet. A single generator script is what keeps them in sync.
Cover, in this order, with a short prefix per app (--n- for Notion):
font: the real platform stack, never a webfont, and measured withrefkit fontin Phase 1 rather than assumed- colour: backgrounds → fills → hairline/border/track → text ramp → accents
- radii, one per component class (
field,card,sheet,tile,pill) - type: composite
font:shorthands, not separate size/weight vars, e.g.--n-t-row: 400 17px/22px var(--n-font) - spacing and geometry constants: gutters, row height, tap target, status bar, sheet top inset
Run sp fetch 7baec1bc-e156-4d3d-97ed-9eada3bc2971 (the luma-ios community project) and
read canvases/luma-ios/ in the folder it prints: a complete run to copy
from, 19 boards (a token board, two evidence boards, 8 screens, 8
references), a four-row layout.json, a committed gen.py, and per-screen
mean deltas of 3.47 to 4.50 levels against the captures.
Run sp fetch 63bcb818-4025-4f93-90de-fa9fd71bc126 (the duolingo-ios
community project) for the second complete run, and the one
to read when the screens are mostly illustration: 58 tokens, 8 screens, 128
pieces of art, and per-screen mean deltas of 1.32 to 2.93, the best
screenshot-sourced numbers in the repo. Every picture on it is a crop of the
capture at a measured box, which is most of why. Its README.md carries the
experiment that settled crop against generate, the stand-in face whose cap
ratio is not SF Pro's, and two defects that produced no error message. Board
00e-art-gen is the generation result, kept as evidence beside the crops it
loses to.
Start from the skeleton rather than a finished board: mkdir -p canvases && cp -r "$KIT/canvases/templates" "$B". Its
gen.py builds the :root block and the evidence table from one TOKENS
list, so a value cannot drift from the evidence behind it and a token cannot
ship without one. Change NAME and the prefix, then replace every
placeholder row with something you measured.
Build the token board as the first generated artboard of the folder (the reference row is already up). It is the contract. When a screen looks wrong later, this is what you check it against.
Once the screens exist, refkit tokens canvases/<slug> enforces the two
invariants this phase rests on: that every board inlines the same :root, and
that nothing references a token that does not exist, in CSS or in the evidence
table. Run it before you call the board done; a --x-scrim-3 in an evidence row
that never existed as a token is invisible to the eye and obvious to the linter.
When the evidence table outgrows the 478 × 980 box, split it onto its own
00b-evidence board rather than trimming it. The evidence is the deliverable.
Phase 3: one generator, N artboards
Write one script that emits every .html file, and commit it with the
boards it produces: canvases/<slug>/gen.py, plus its asset JSON,
resolving paths relative to __file__ so
python3 canvases/<slug>/gen.py regenerates the folder in place
($KIT/canvases/templates/gen.py is the skeleton,
$(sp fetch 7baec1bc-e156-4d3d-97ed-9eada3bc2971)/canvases/luma-ios/gen.py a
finished one, and
$(sp fetch 63bcb818-4025-4f93-90de-fa9fd71bc126)/canvases/duolingo-ios/gen.py
a finished one that also cuts and
places its own artwork from a crops.json). Do not hand-edit the
artboards afterwards; edit the generator and re-run. That is what keeps
eight files consistent through a dozen correction passes, and it only
outlives the session if the generator is in the repo; a scratch-dir
generator dies with the session and leaves the boards unmaintainable under
this skill's own rules. A board with no committed generator is
unfinished.
TOKENS = ":root{...}" # from Phase 2, one source of truth
BASE = ".phone{width:393px;height:852px;...}"
SB = '<div class="sb">...</div>' # status bar, 54px
def page(title, extra_css, body): ... # TOKENS + BASE + extra_css + body
def write(name, html): ...
Hard constraints from the canvas renderer (also in sp-canvas's
references/layout.md):
-
Fully self-contained. The iframe is
sandbox="". No external CSS, JS, fonts or images. Every image is adata:URI; icons are inline SVG. Keep each icon asassets/icons/<name>.svgand inline it through a helper ingen.py: the canvas's inspector names an inline<svg>from that folder by its geometry, and hands it back as a vector asset. A glyph written as a literal ingen.pygets no name. -
No page ground. The canvas releases the frame's own opaque white background, so a board shows the canvas through wherever it paints nothing. The shared
body{}rule ingen.pyis where this gets broken. Leave it with nobackground, astemplates/gen.pyships it, and let.phonepaint its own, so a screen board floats on the canvas instead of sitting on a white card. Document boards, the token sheet and the evidence sheets, are the exception. They are a page rather than a device, and their black text needs a ground. -
Artboard box is 478 × 980. Overflow clips silently. Do not check this by eye;
refkit shoot ... --check-overflowasks the layout engine and exits non-zero with the exact px, so a clipped board fails in Phase 3 rather than turning up in Phase 4. It also flags elements clipped inside their own containers; mark a deliberate clip (a fade, a masked hero) withdata-clip-ok. -
Phone frame is 393 × 852 pt at 1pt = 1px: 54px status bar, 125 × 36 Dynamic Island, 139 × 5 home indicator.
-
The iOS status bar comes from the template, never from the capture. Any iPhone screen with a status bar uses the one
templates/gen.pyships finished:PHONE's.sbrules, theSB_ICONScellular/Wi-Fi/battery SVGs andstatusbar()with its9:41clock and plain island. Copy those four across unchanged and callstatusbar()as it is. Measuring four signal bars again costs an hour and lands within a pixel of what is already there.The one thing that changes is the frame. On a wider frame, move the
SB_ICONSgroup by the width delta so the battery keeps its right inset.Nothing in the capture's own status bar is carried over, whether drawn or cropped. That covers its clock and any extra glyph: a mute bell, a Focus or person badge, a location arrow. It also covers an expanded island or Live Activity, an iOS 26 filled battery, a different charge level and no-service bars. The status bar is shared chrome across every board in the repo, and a per-capture copy is exactly the drift the rule is there to stop. The diff pays a few levels in the top 54 pt for it, and that is expected.
-
No board background on a screen artboard. Give
bodynobackgroundat all, so the phone floats on the canvas ground and its drop shadow lands on whatever the board is placed over. A cream or grey field behind the phone paints one opaque rectangle per artboard, and a row of those reads as jarring next to its neighbours. The exceptions are the full-bleed sheets, the token board and the evidence boards, which are their background and keep it. -
Avoid SF Symbols private-use glyphs; they render as tofu without SF Pro installed. Inline the SVG, or embed a rasterized symbol as a
data:URI. -
Icon size is not a calibration loop. Set each SVG's
viewBoxto the measured ink box, in pt, and give the span the same numbers; scale is then 1:1 with nothing to converge. The defaultpreserveAspectRatio(xMidYMid meet) means that when the span's aspect ratio disagrees with the viewBox's, only one axis binds, and a size loop silently converges on the wrong axis and stops improving.
Copy is part of the replica. Transcribe the reference's strings exactly,
including where each line wraps. A title that breaks after "iPhone and"
instead of "iPhone and AirPods" is a real defect. Force it with explicit
widths, <br>,   or <wbr> rather than hoping the browser agrees.
When a substituted face fights a measured width, the wrap wins. A stand-in that sets wider pushes a string onto a second line inside a container the capture shows holding one, and two measured rules collide. Widen the container and record why; never shrink the type to make a measured width hold. Claude's user bubble measures 302.6 and ships at 316.
Artwork: crop only the picture, rebuild the interface on it
A screen that is mostly illustration is not mostly work, but a screenshot pasted into the frame is not a replica either. Sort the pixels three ways before cutting anything:
- Interface is rebuilt, however much it looks like a picture: type,
buttons, pills, badges, chips, tab glyphs, sheet marks, a keyboard, an
illustration made of flat fills. Use HTML and CSS, and trace each glyph from
the capture into an SVG whose viewBox is its ink box in page points — trace a
padded crop and let the tracer measure that box, because a crop cut to the
glyph traces its apexes flat. Use the
system font's own outline where one exists, such as SF's Apple logo at
U+F8FF. A crop of interface scores well and is worthless, because nothing
in it can be edited, reflowed or reused. A trace is not the finish for a
glyph the board draws large. It is a polygon of hundreds of linetos whose
edges show facets at 2× while every delta reads 0.
references/glyphs.mdredraws such a glyph as the primitives or the curves it was designed as. - An app icon or third-party logo is the original file, never a crop:
the iTunes lookup API's 1024 px artwork for an App Store app, and
references/brand-marks.mdfor the rest. - Only a picture is cropped, meaning photography, illustration and
editorial art that the capture is the only source for. It is keyed by id in
a
crops.jsonthe generator reads, cut toassets/art/<id>.pngand placed back by anart()helper at the same numbers. The asset then cannot drift from where it was measured, and its pixels are the reference's own.
When interface is set on a picture, such as an app lockup on a hero or a
headline on a card, crop the picture and erase the interface out of it.
List it in the crop's erase: a box for an icon or pill, and a box plus a
threshold for type. cut() inpaints those pixels from the ones around them,
and the board draws the interface live on top. End the crop where the picture
itself ends (its fade to the ground), not at the first line of type on it. If
the artwork is published anywhere else, give the crop a registered guide so
the fill under wide type is the artwork's own texture. See assets.md. Otherwise the board carries
every word twice, once in the picture and once in the type, and the copy in
the picture cannot change. The only thing cut whole is a sliver too small to
identify, such as a 10 pt peek of the next card at a scroll edge. There is
nothing to rebuild it from. Rebuilding costs levels a crop would not.
apple-app-store went from 2.32-8.00 to 3.62-8.02 when its interface came out
of the crops, and that is the trade to make.
For a picture, the rule and the arithmetic behind it: a crop scores 0 by
construction, and the same crop redrawn by gpt-image-2 scores 38.53, the same
character with a different head-to-body ratio and the props moved.
So generate only pixels the capture does not contain, name those assets
in the folder README, and give each one a probe like any other measurement.
When you do have to generate, the thing that decides the result is the
input's layout, not the prompt: pack the assets into a grid, each in its
own cell at the size and position it must come back at, and the model
upscales in place instead of composing. That took this repo's set from 18.41
to 3.96 mean delta, with every cell returning at scale 0.99-1.00. Density
is free (77 assets in one call beat 6), but the asset's native size in the
capture is not: under about 128px, colour does not survive the redraw, so
small icons stay CSS or inline SVG. artgen runs the whole loop,
including keying the cells back out, solving the fit and scoring each asset
against the crop it came from.
references/assets.md has the decision table, the
crops.json/cut()/art() shape, the erase list and its traps, and why
assets/art/ is committed while assets/refs/ is not;
references/generating.md
is the generation procedure end to end, with the key-colour, alpha-ramp and
fit-sign traps that each cost a run;
references/glyphs.md redraws the traced glyphs a
board draws large, by construction or with refkit refit.
Model the line box once, then place by ink
A screen carrying real prose is a dozen placements of the same few numbers,
so measure the block model once with bands and let the generator solve the
rest: the line box per level, the margin between each pair of block types,
and the offset from a line box's top to the ink inside it. Then write the
builder to take an ink top and solve back to the box, so every call site
is a number read straight off the grid rather than one derived by hand.
K = {"sans": 0.2708, "serif": 0.2240} # SF Pro, Georgia: cap top within the line
def ct(ink, fs, lh, face="sans"): # the box top that puts ink where you want it
return ink - ((lh - fs) / 2 + K[face] * fs)
Claude's answer column is one such model: h1 29/36.3, h2 25/31, body
17.8/25.5, margins of 8.4 generic, 5.7 into a heading, 10.3 out of one, 8.2
between list items. Measured once, it placed five long screens to within
0.5pt. The alternative is a hand-tuned top per paragraph, which drifts the
moment any string above it changes.
Never let content you invented carry a measured element into position. Where a composer, a fade or a scroll edge hides part of a block, the hidden span is not on the capture and cannot be transcribed. Filling it with plausible prose so that the flow pushes the visible remainder down makes that remainder's position an invention, and it will read as measurement to everyone after you. Place what you can see as its own absolutely-positioned block on its own measured ink top, and treat the gap as a line count or as nothing at all. Say so in the folder README.
layout.json already holds the reference row (Phase 5 builds it first). Add
the Foundations row and the numbered-screens row as the boards land.
Phase 4: verify by rendering, not by reading
Render at the capture's own scale and cut the screen out of the frame, so your render and the reference share one pixel grid, then replay Phase 1's probes against the renders before anything else:
refkit shoot "$B"/[01]*.html \
-o "$B/scratch/mine" --scale 3 --crop-phone --check-overflow
refkit batch probes.json --against "$B/scratch/mine"
refkit diff "$B/scratch/mine/07-models-sheet.png" "$B/assets/refs/cp7.png" \
--pt 3 -o "$B/scratch/d07.png" --regions regions.json
--scale takes an integer, and your capture's scale is not one. A
2.2417 px/pt capture against a 3× render is a shape mismatch and diff
refuses it. Shoot at the nearest integer and downscale each render to the
capture's own pixel size (Pillow, LANCZOS) before diffing. Put the whole
chain, regenerate to shoot to downscale to diff all N, in one scratch script
on the first iteration; you will run it thirty times.
batch re-runs every measurement that justified a token, reference against
render, in one table. crops probes.json --against mine -o crops/ writes a
paired NEAREST-upscaled crop per probe. Size is a numbers problem, shape is
a picture problem: the crops are what find "too small, too small, and
rotated 10 degrees" while the numbers read fine.
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 77
- Forks
- 3
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
sp-clone-prototype- Source
- github.com/resciencelab/super-prototyping