Weekly review
SkillProductivityWeekly review, weekly scorecard, how was my week, week in review, weekly rollup, Sunday review, Friday review, what did I get done this week, next week's top three, weekly retrospective, weekly check-in. Composes the week's meetings and hours, commitments closed versus dropped versus open, leads captured, money findings, and content shipped into one scorecard by reading the sibling routines' own reports rather than re-deriving them. Leads with the multi-week trend rather than this week's figures, carries provenance on every number, is willing to state plainly that the week was poor, and selects next week's top three on consequence ahead of urgency with an escalate-or-drop rule for carried items. Runs as a weekly routine, or on demand. Requires the Littlebird MCP.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Weekly review skill
What this skill tells your AI
The instructions your AI receives, as published by legioncodeinc/vibe-coding-tools in src/quarantine/skills/keep/weekly-review/SKILL.md and read by ahel’s review.
Purpose
One scorecard per week covering the whole operation: meetings held and hours spent in them, commitments closed against dropped against still open, leads captured and what happened to them, money findings, content shipped, what moved per project, and next week's top three with the reasoning shown.
This is the composition skill of the marketplace, and that is its defining property.
Nearly everything in the scorecard was already produced by a sibling routine. So the primary
retrieval of this skill is LB_INTERNAL_LIST_ROUTINES plus LB_INTERNAL_GET_ROUTINE_REPORTS
across the user's other routines, reading their weekly output. Exactly one section is
retrieved fresh every run: meetings and hours.
Read, do not re-derive. Re-deriving is slower and more expensive, and worse than both, it produces a number that disagrees with the sibling's own published number. The user then has to reconcile two versions of their own week. For a scorecard that is the worst available outcome, because one figure the reader cannot trust undermines every other figure on the surface [references/research/distilled-weekly-review-design.md, section 7].
Two design properties matter more than the rest.
- The trend is the product, not the snapshot. A single week's numbers mean almost
nothing. The routine reads twelve of its own past reports and the scorecard leads with
direction.
references/trend-construction.md. - Honest scorekeeping, which means being willing to report a bad week. A review that
always finds something positive is worthless within a month and it is the single most
likely way this skill fails.
references/honest-scorekeeping.md.
What the evidence actually supports. Not the weekly review, which this archive found no controlled evidence for at all. Monitoring: across 138 studies and 19,951 participants, monitoring goal progress moved goal attainment by d+ = 0.40, mediated by monitoring frequency, with larger effects when the information was physically recorded and when outcomes were reported [references/research/distilled-weekly-review-design.md, section 1]. A persistent routine report is a physical record by construction. That is the whole warrant for this skill, and it is a moderate effect, not a transformation.
Capability gate
This skill requires the Littlebird MCP on a Power or Pro plan.
Before anything else:
- List the tools actually available in this session. Do not assume tool names. Confirm
that the routine tools, the meeting tools, and
search_user_contextare present under their real names. - If the Littlebird MCP is not connected, stop and tell the user: "This skill needs the Littlebird MCP connected on a Power or Pro plan. Connect it at https://support.littlebird.ai/docs/mcp/ and run this again."
- If the routine tools are missing, this skill is severely degraded and says so. Without
LB_INTERNAL_GET_ROUTINE_REPORTSthere is no sibling rollup and no series, which is most of the value. Run on-demand mode against fallbacks only, print one line at the top saying the scorecard is running without any rollup or trend, and do not present it as a weekly review. - Before creating the routine, call
LB_INTERNAL_GET_SUBSCRIPTION_STATUSto confirm the plan allows another routine. Routine count is plan-limited, so if the account is at its limit, name which existing routine should be replaced rather than proposing an addition [references/littlebird-mcp-reference.md].
Tool mechanics, parameters, and return shapes: references/littlebird-mcp-reference.md.
Littlebird MCP calls used
| Tool | Used for |
|---|---|
LB_INTERNAL_LIST_ROUTINES | Discovering which sibling routines exist, their schedules, their latest report dates, and whether they are paused. The primary retrieval. |
LB_INTERNAL_GET_ROUTINE_REPORTS | Two things, both mandatory: twelve of this routine's own past reports for the series, and two reports from each matched sibling for the rollup |
LB_INTERNAL_LIST_MEETINGS | Meetings held and hours in meetings for the window. The one section always retrieved fresh, and the only genuinely measurable number in the scorecard |
LB_INTERNAL_GET_MEETING | Fallback only. The ## Action Items and ## For You sections when no commitment sibling reported |
search_user_context | Fallback only. Money, leads and content sections when their siblings are absent, stale or paused |
LB_INTERNAL_GET_SUBSCRIPTION_STATUS | The plan and routine-slot check before creating the routine |
LB_INTERNAL_CREATE_ROUTINE | Creating the weekly routine, from an interactive session only |
LB_INTERNAL_GET_ROUTINE_CONFIG and LB_INTERNAL_UPDATE_ROUTINE | Changing the routine later. Read the config first, because prompt and schedule each replace the whole field |
Never used by this skill: LB_INTERNAL_GET_MEETING_TRANSCRIPT and
LB_INTERNAL_SEARCH_MEETINGS. A weekly scorecard has no budget for transcript reading, and a
topic search over meetings is a deep-run instrument that belongs to the siblings.
Trigger
Trigger phrases: weekly review, weekly scorecard, how was my week, week in review, weekly rollup, Sunday review, Friday review, what did I get done this week, what should I focus on next week, next week's top three, weekly retrospective, weekly check-in, set up my weekly review.
Do not trigger for: today's plan (that is daily-brief), a full commitment ledger (that is
commitment-tracker), a per-client risk view (that is client-health-radar), a vendor spend
audit (that is money-leak-auditor), or an audit of whether the routines themselves are
healthy (that is routine-architect).
Routine cadence
Weekly, plus on demand.
The timing position, taken deliberately: Friday late afternoon by default, roughly 16:30 local. Sunday evening is fully supported, offered every time, and is the second choice.
The practice literature offers Friday afternoon, Sunday evening and Monday morning with no evidence behind any of them, then says consistency matters more than the choice [references/research/distilled-weekly-review-design.md, section 9]. So the decision comes from the recovery literature instead.
Boundary management is the strongest lever on weekend detachment by a wide margin, d = 0.65 for interventions with a boundary-management component against d = 0.25 without [references/research/distilled-weekly-review-design.md, section 9]. A scheduled, notified Sunday-evening work scorecard is boundary management run backwards. And the recovery paradox bites this skill specifically: perceiving that one has performed well predicts better evening detachment, so an honest report that must sometimes state the week was poor would deliver that to the reader least able to detach afterwards [references/research/distilled-weekly-review-design.md, section 9].
The strongest objection to Friday, that it is inside the work week and Friday judgment is tired, mostly dissolves, because the generator is a routine and the machine has no Friday afternoon. It runs at 16:30 against a week that is effectively complete, and the human reads whenever they choose, including Monday morning.
Sunday evening keeps a real case: not detaching predicts positive affect when the thinking is
problem-solving rather than rumination, and very high detachment may itself undermine
performance [references/research/distilled-weekly-review-design.md, section 9]. Offer both
with the tradeoff in plain terms, set what the user picks, and do not argue past one answer.
Full argument in references/honest-scorekeeping.md, section 6.
Two modes.
| Mode | Trigger | Output |
|---|---|---|
| Routine (primary) | Scheduled weekly | One routine report per run, at or under 450 words, ending in the SERIES line |
| On demand (secondary) | User asks | The same scorecard plus an appendix, written to a file |
Process
Step 1: read your own past reports
Mandatory, first, before any other retrieval. LB_INTERNAL_GET_ROUTINE_REPORTS on this
routine with limit: 12.
Parse the SERIES line at the end of each report rather than re-reading the prose. Build the series for every field, and build the consecutive-week count for every item that has appeared in a top three.
Twelve, because the stricter published shift rule is defined for a series of 12 to 22 points,
and because twelve weeks is a quarter, which is a window the reader can check against memory
[references/research/distilled-weekly-review-design.md, section 6]. The series line format and
the parsing rules: references/trend-construction.md, section 2.
A routine prompt that does not instruct the model to read its own previous reports will repeat itself indefinitely [references/littlebird-mcp-reference.md].
Step 2: roll up the siblings, do not re-derive them
LB_INTERNAL_LIST_ROUTINES with limit: 25, then LB_INTERNAL_GET_ROUTINE_REPORTS with
limit: 2 on each matched sibling. Match on the substance of the title, not on an exact
string.
This is the primary retrieval of the skill, not a preliminary step. In a well-populated account it produces five of the six scorecard sections.
The section-by-section mapping, the weekly freshness gate, the exact lines to print for a
stale, paused, absent or genuinely-empty sibling, the per-section fallback queries with their
budget cap, and the provenance marks: references/rollup-and-fallbacks.md.
The general rollup pattern this guide builds on lives in the daily-brief skill, in its
reference guide named rollup-composition, and is not restated here. This skill's guide
stands alone where daily-brief is not installed.
Step 3: retrieve meetings and hours, always
The one fresh retrieval. LB_INTERNAL_LIST_MEETINGS across the window. This is genuinely
measurable and it is one of the few honest numbers in the whole scorecard, which is exactly
why it must carry its bound: scheduled duration is not attendance.
Step 4: run the fallbacks, only for uncovered sections
At most five calls. Priority order when the budget binds: commitments, money, leads, content,
projects. Every fallback result carries its reduced-check line.
references/rollup-and-fallbacks.md, section 5.
Step 5: build the series and decide what the report may say
Append this week's values, then apply the length table. One point licenses no direction claim at all. Two licenses a one-week change and the words "trend", "improving" and "momentum" are banned. Three to four license an early indication, in the source's own hedged wording. Five consecutive rising or falling license the word trend. Twelve or more license the word shift.
A rule is never evaluated across a gap in the series. Closing up an na to reach five
consecutive points is manufacturing a trend.
references/trend-construction.md, sections 3 and 8.
Step 6: order the scorecard by what the series says
Direction first, absolute figures second. The section whose series is doing the most interesting thing leads: shift, then trend, then astronomical point, then a crossed threshold, then template order. A report that opens with this week's counts is the measurement-instead-of-decision failure that gets scorecards abandoned [references/research/distilled-weekly-review-design.md, section 7].
Step 7: apply the honesty gates
Print the poor-week block when its triggers fire, plainly, at the top, with no cushioning clause. Refuse both manufactured wins and manufactured crisis. Keep every sentence at the level of the work rather than the level of the person, because that is the axis the feedback evidence actually supports: over a third of measured feedback effects made performance worse, and the mechanism is attention moving from the task to the self [references/research/distilled-weekly-review-design.md, section 4].
The six banned win-manufacturing moves, the five banned crisis-manufacturing moves, the exact
shape of the poor-week block, and the three self-diagnosis signatures:
references/honest-scorekeeping.md.
Step 8: select the top three
Score every candidate on consequence at weight 3, urgency at weight 2, carry at weight 2. Filter spurious urgency to zero before scoring rather than penalizing it during. An item may reach the top three on consequence alone, and that is the intended behavior.
Any item in the top three for three consecutive weeks gets the escalate-or-drop block. At four it is dropped by default.
Candidate pool, the scoring tables, the five tie-breaks, the carried-item block, the honest
statement of why three, the four-line output shape with its mandatory Beat line, and the two
no-pick cases: references/top-three-selection.md.
Step 9: count, cut, and write the SERIES line
At or under 450 words. If over, delete whole items lowest-ranked first. Cut items, never cut evidence. Do not get under the ceiling by dropping receipts, stripping provenance marks, merging findings, or removing a Beat line.
Then write the SERIES line as the last line of the report, with na for anything unmeasured
and ~ for anything from a fallback.
Retrieval brief
The actual calls. Substitute real dates; never leave a placeholder in a live call. The window is the seven days ending on the run date.
Own history, once per run
LB_INTERNAL_GET_ROUTINE_REPORTS
routine_id: [this routine's id]
limit: 12
Sibling discovery, once per run
LB_INTERNAL_LIST_ROUTINES
limit: 25
Sibling reports, once per matched sibling
LB_INTERNAL_GET_ROUTINE_REPORTS
routine_id: [sibling id]
limit: 2
Meetings and hours, always, once per run
LB_INTERNAL_LIST_MEETINGS
start_date: [window start]
end_date: [window end]
limit: 60
Returns both recorded meetings and unrecorded calendar events; only recorded ones carry an id [references/littlebird-mcp-reference.md]. Count both, and report the split, because the unrecorded ones are calendar entries the user may or may not have attended. Hours come from scheduled duration, which is a bound and not attendance, and the number carries that bound.
Commitments fallback, only when no commitment sibling reported
LB_INTERNAL_LIST_MEETINGS
start_date: [window start minus 21 days]
end_date: [window end]
limit: 40
then, on at most eight recorded entries that carry an id:
LB_INTERNAL_GET_MEETING
meeting_id: [id]
Read only ## Action Items and ## For You. Those sections already carry owner attribution
[references/littlebird-mcp-reference.md].
Money fallback, only when no money sibling reported
search_user_context
search_queries: ["subscription renewal charge", "invoice overdue payment", "annual plan renews on",
"payment failed card declined", "your card will be charged"]
standalone_query: "Billing, renewal, invoice and payment notices that appeared on screen this week"
date_range: {"start": "[window start]", "end": "[window end]"}
filters: {"data_source": "snapshots"}
Leads fallback, only when no lead sibling reported
search_user_context
search_queries_messages: ["interested in", "send me the details", "how much is it", "can we talk",
"dropped you a DM", "want to learn more"]
standalone_query: "New people who expressed interest in what I sell during this week"
date_range: {"start": "[window start]", "end": "[window end]"}
filters: {"data_source": "messages"}
Content fallback, only when no content sibling reported
search_user_context
search_queries: ["published post", "just posted", "newsletter sent", "video uploaded", "went live"]
standalone_query: "Things I actually published or sent this week, as opposed to drafted"
date_range: {"start": "[window start]", "end": "[window end]"}
filters: {"data_source": "summaries"}
Prefer several narrow parallel queries over one broad one, both for relevance and to avoid the oversized-result file dump [references/littlebird-mcp-reference.md].
Evidence standards
Every line obeys references/evidence-standards.md. The rules that bite hardest here:
- Provenance on every number. Which sibling report and its date, or which retrieval and
its window. Plus exactly one mark:
(exact),(bounded: reason), or(reduced check). A number with no mark does not go in the scorecard. - Observed, inferred, external, unknown. Each line is exactly one, visibly. The top three is inference by construction and carries the observations it rests on.
- Absence is not a negative finding. "No evidence in the record since 2026-08-05", never "they did not do it". A section with no sibling is not a zero, and the scorecard prints the reason instead of a number [references/rollup-and-fallbacks.md, section 4].
- Confidence ratings. A Low-confidence claim never enters the top three and never gets an urgency score above zero.
- Attribution guardrail. Capture shows what the user was viewing, not what they wrote. A composer window is not published content.
- Partial rosters are reported as partial. Lead counts from message capture are floors, because platform UIs collapse lists.
- Rolled-up claims keep the sibling's hedge and the sibling's confidence. Never restate a sibling more confidently than the sibling did. A scorecard compresses, and compression is where a hedge gets dropped.
- Relevance scores. Anything scored 3 is a maybe. Do not build a scorecard number on a single 3-scored result [references/littlebird-mcp-reference.md].
- Sensitive categories stay out. Health, financial detail beyond business figures, legal history, family circumstances, protected characteristics, and precise home location, even where the capture contains them.
- Raw capture never ships. Process in temp space, produce the scorecard, delete the raw.
- Confirm before encoding. On-demand mode confirms with
AskUserQuestionbefore recording a durable fact about a person or a figure. Routine mode cannot ask, so routine mode does not encode durable facts; it reports with hedges and marks.
Draft never send
This skill drafts and holds. Nothing is sent, posted, published, or written into a third-party
system without the user approving the actual final text through AskUserQuestion. Approving a
plan is not approving the words. This applies even where a Gmail, Slack or CRM connector is
connected in the session.
The weekly review produces no outbound text at all in normal operation. If the user asks it to
chase something the scorecard surfaced, hand off: commitment-tracker owns nudges,
invoice-chaser owns receivables chasing, renewal-sentinel owns cancellations.
If a connector is needed, list the available tools first and degrade gracefully when it is absent: produce a copy-paste block rather than assuming a connector exists.
Empty retrieval
Five distinct empty cases. None of them fabricate.
No sibling routines exist at all. Run every fallback within the five-call cap, print the
no-sibling line for each uncovered section, and add one line at the top: Running without any rollup. Every figure below is a reduced check. The siblings that would produce real figures are named per section. Then offer to set up the two highest-value siblings for this user.
A sibling is stale or paused. Print the stale line or the paused line, run the fallback, and never omit the section [references/rollup-and-fallbacks.md, section 4]. A missing section reads as a zero and a zero is a claim.
A section genuinely has no items and the sibling looked. Print the zero with the sibling's report date. This is the only case where a bare zero is correct, and the distinction between "the sibling found none" and "nobody looked" is the point of the whole section.
A quiet week: everything retrieved, nothing moved. Print the short form: the series lines with their flat readings, one sentence saying the week was flat, and the top three if any item qualifies. Do not lower the bar to fill a section, do not promote a minor item, and do not manufacture urgency to justify the report existing. Real quarters contain quiet weeks [references/honest-scorekeeping.md, section 3].
Everything came back empty. Report the gap and stop:
No Littlebird data retrieved for this window. No routines found and no capture returned.
Nothing to review. This usually means capture was off or the account has no recent activity.
Never pad from training data, never reason from what would probably be there, never substitute plausible examples [references/evidence-standards.md, rule 9].
Output
Routine mode produces one Littlebird routine report per run, titled
Weekly review, week ending [Month D, YYYY], at or under 450 words, in this shape.
| Part | Cap | Contents |
|---|---|---|
| Lead | 3 lines | The section whose series is doing the most interesting thing, direction first. The poor-week block goes here when triggered. |
| Meetings | 1 line | Held, split recorded against calendar-only, hours with its bound, series |
| Commitments | 3 lines | Closed of total as a rate, dropped in full by name, still open, series |
| Leads | 1 line | Captured, how many have a next step recorded, series |
| Money | 3 lines | Leaks found, renewals inside 14 days, receivables outstanding, each with provenance |
| Content | 1 line | Shipped, by name and date. Drafts excluded and said to be excluded |
| Moved and did not move | 4 lines | Per active project, state changes only |
| Next week's top three | 12 lines | Three items, four lines each, Beat line mandatory |
| Carried block | Only at 3 weeks | Escalate or drop, per top-three-selection.md |
| Selection note | Only when 2 or more carried | The self-diagnosis line |
| Footer | 2 lines | Retrieval date, which sibling reports were rolled up and their dates, then the SERIES line |
The 450-word ceiling is a design element, not tidiness. The practice literature names arduousness as the reason the habit dies: "the longer and more arduous your review is, the less likely you'll be to maintain the habit" [references/research/distilled-weekly-review-design.md, section 2]. And the recovery literature's distinction is between bounded problem-solving thinking, which is benign, and open-ended rumination, which is not [references/research/distilled-weekly-review-design.md, section 9]. A short report ending in three decisions is the first thing. An open-ended reflective essay is the second. A stated ceiling does not produce a ceiling, so the ceiling appears four ways: per-section caps, the ordering rule, an explicit count-and-cut step, and a ban on getting under it by cutting evidence.
On-demand mode produces a file at weekly-review-[YYYY-MM-DD].md in the working
directory: the identical scorecard, plus an appendix holding the full twelve-point series for
every field, the sibling report dates and staleness state for each section, the full candidate
pool for the top three with every score shown, and the items excluded with the reason each was
excluded. Nothing in the appendix is required reading. State the path to the user when done.
Both modes end with the retrieval date, the rolled-up sibling reports and their dates, and the SERIES line, so the next run can rebuild the series from one line per report.
Guardrail
The specific risk this skill carries is authority laundering: a number acquires authority by being printed in a scorecard, regardless of where it came from.
Three failure paths, all specific to this skill's shape.
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 83
- Forks
- 37
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
weekly-review-legioncodeinc- Source
- github.com/legioncodeinc/vibe-coding-tools