ACM MM Related Work

SkillMedia

Build or audit the related-work section of an ACM Multimedia (ACM MM) paper. Once added, your AI can cover the multimedia literature across vision, audio and speech, language, HCI and QoE, and systems, keep citations double-blind, and account for concurrent work appearing quickly on arXiv.

Available today. Use it from your connected AI after setup.

After adding it, share your paper draft or related-work notes and ask your AI to build or audit that section of your ACM MM paper.

Then ask your AI: use the ACM MM Related Work skill

What your AI can do with it

  • Draft a related-work section for an ACM MM paper
  • Audit an existing related-work section for coverage and accuracy
  • Cover literature spanning vision, audio/speech, language, HCI/QoE, and systems
  • Track concurrent work appearing on arXiv at speed
  • Keep citations double-blind for review
  • Verify papers cited as ACM MM work

What this skill tells your AI

The instructions your AI receives, as published by brycewang-stanford/awesome-journal-skills in ACM-MM-Skills/skills/acmmm-related-work/SKILL.md and read by ahel’s review.

Use this to position an ACM Multimedia paper against a literature that is unusually spread out: a cross-modal contribution touches several single-modality communities plus the multimedia venues that combine them.

The five multimedia shelves

A strong ACM MM related-work section shows command of all shelves the contribution touches, not just the author's home community:

  • Vision — the visual side (CVPR/ICCV/ECCV) your method builds on or competes with.
  • Audio/speech and music — the acoustic side (ICASSP, INTERSPEECH, ISMIR) if a stream is audio.
  • Language — the text side (ACL/EMNLP) if captions, transcripts, or descriptions matter.
  • HCI / QoE / human-centric — perception, engagement, and interaction (CHI, QoMEX) when the claim is subjective.
  • Multimedia proper — ACM MM, ICMR, MMSys, and TOMM, where these threads are combined.

Coverage vs. venue-hygiene table

TaskWhat to doFailure it prevents
Cover each modality you useCite the current best single-modality work per stream"You ignored the vision literature"
Cite the fusion lineageTrace the cross-modal line your method extends"No delta over existing fusion"
Verify the venue stringConfirm each "ACM MM" cite on dblp conf/mmMisattributing an ICMR/CVPR paper to ACM MM
Handle concurrencyNote contemporaneous arXiv work honestly"You missed / overclaimed novelty vs. X"
Keep it blindCite your own prior work in third personDouble-blind violation

Positioning, not listing

Do not enumerate. For each closest neighbor, state in one clause what it did and in one clause what your paper adds — and make sure the delta is cross-modal, since "we swap a better encoder" is a single-modality delta a reviewer will discount.

[Neighbor] Late-fusion audio-visual highlight scoring (VenueYear).
[Their move] Average per-modality scores.
[Our delta] Score the timing DISAGREEMENT between streams — a signal averaging cannot represent.

Concurrency under arXiv speed

Multimedia sub-areas move fast and preprint heavily. Acknowledge genuinely concurrent work (roughly same-window preprints) as concurrent rather than prior, do not claim to beat a method you did not run, and do not silently drop a close preprint — a reviewer who knows it reads the omission as evasion.

Double-blind citation hygiene

  • Cite your own earlier papers in the third person ("Prior work [X] showed..."), never "our previous paper."
  • Avoid anonymity leaks through a dataset, system name, or repository that only your group uses.
  • The single-blind tracks (Reproducibility, Open Source Software, Dataset) relax this — but the main track and Brave New Ideas do not.

Building the section in passes

Pass 1: list the modalities and sub-fields your contribution touches (the shelves).
Pass 2: for each shelf, cite the current strongest 2-3 works you build on or beat.
Pass 3: trace the FUSION lineage — the cross-modal line your method extends — as its own thread.
Pass 4: add same-window arXiv work as concurrent, not prior.
Pass 5: verify every "ACM MM" cite on dblp; convert your self-cites to third person.

A reader should finish the section able to name which prior fusion approach you improve on and why the single-modality shelves are covered but not the whole story.

The multimedia-specific omission risk

Because a cross-modal paper sits between communities, the dangerous omission is usually the other community's closest work — a vision-trained author who misses the audio or IR paper that already did half the job. Before submitting, ask a question from each shelf's perspective: "what would an audio reviewer, an IR reviewer, and a systems reviewer each say I missed?" That triage catches the omissions that sink cross-area papers.

Venue-verification pass

Before submission, spot-check every citation that claims an ACM MM placement against dblp's conf/mm edition record. The common traps: ICMR papers cited as ACM MM, TOMM journal articles cited as the conference, and CVPR/ICCV vision papers cited as multimedia. Fix the venue string or the claim that rests on it.

Output format

[Shelf coverage] vision/audio/language/HCI-QoE/multimedia — <covered / gaps>
[Fusion lineage] traced / missing
[Deltas] cross-modal and specific / single-modality or vague: <list>
[Concurrency] handled / risky omissions: <list>
[Blindness] clean / leaks: <list>
[Venue hygiene] verified / suspect cites: <list>

Signals

GitHub stars
1k
Forks
146
Last commit
Aug 2026
Advanced
Catalog kind
skill
Gateway key
acmmm-related-work
Source
github.com/brycewang-stanford/awesome-journal-skills