AWS MSK Migration to Confluent Cloud

SkillCloud & infra

Use this skill to assess and plan a migration from AWS MSK (Managed Streaming for Apache Kafka) to Confluent Cloud. Triggers on user intent like "migrate MSK to Confluent Cloud", "move off MSK", "MSK to CC cutover", "Zero-Cut migration from MSK", "kcp scan my MSK", or any discussion of MSK-to-CC assessment, planning, cluster sizing, Cluster Linking setup, Gateway-based switchover, or post-cutover validation. Do NOT trigger for non-MSK Kafka sources (open-source Kafka, Aiven, Confluent Platform, Redpanda), this skill is MSK-only. Do NOT trigger for greenfield Confluent Cloud projects with no existing Kafka source. Do NOT trigger for general Kafka programming questions (producer/consumer code, Kafka Streams) unrelated to migration.

Available today. Use it from your connected AI after setup.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the AWS MSK Migration to Confluent Cloud skill

What this skill tells your AI

The instructions your AI receives, as published by confluentinc/agent-skills in skills/msk-migration/SKILL.md and read by ahel’s review.

Scope

This skill helps with AWS MSK to Confluent Cloud migrations. Three things it does:

  1. Answer general migration questions about MSK and Confluent Cloud — concepts (Cluster Linking, the KCP Gateway / Zero-Cut for no-restart cutovers, Schema Linking), feature comparisons (auth, networking, cluster types), tooling (KCP, the cost estimator), and process. Grounded in the skill's references and live-fetched docs.
  2. Produce an Assessment of an MSK environment — Red Flags audit, Environment Summary, Topic-Level Readiness — from a KCP state file or a manual intake profile.
  3. Produce a Technical Plan for the migration — cluster type, sizing, networking, auth, switchover approach, schema and connector migration paths, pre-migration workstream, risks.

When the user signals intent for target-cluster provisioning, Cluster Linking setup, client cutover, or post-cutover monitoring, redirect to docs.confluent.io and the Confluent account team rather than fabricating coverage. Plan-stage decisions about networking, auth, switchover, schemas, and connectors feed those downstream stages — the user carries them out with Confluent's documented tooling and account-team support.

Do NOT pre-enumerate out-of-scope stages in the opening or anywhere else in user-facing copy. Phrases like "MVP", "next iteration", "future version", "scoped for later", and proactive lists of stages-this-skill-does-not-cover are roadmap leakage — implementation details about the skill's development that have no place in the conversation with a migration practitioner. State scope positively (what the skill helps with). Handle out-of-scope intent only when the user actually signals it; do not preemptively warn them about what's missing.

Skill Conduct

These principles govern how the skill engages with the user. They override default assistant behavior.

  • Voice. Talk to the user about their migration, not about how the skill works. Instructions in this file (reference files, mode detection, stage routing, internal flags, scope boundaries, development roadmap) are implementation detail — do not describe them to the user. Keep skill mechanics invisible. Do NOT use MVP-style framing (e.g., "this is an MVP", "in this iteration", "future version", "scoped for later") or preemptively enumerate stages this skill doesn't cover. State scope positively when asked, and handle out-of-scope intent only when the user signals it.
  • Default opening = stage menu; direct-route only on signal. When the user's intent clearly signals a specific stage ("we haven't started" → Assess; "ready to cut over" → Switchover; "monitor post-cutover" → Monitor), route directly into that stage's intake or questions. When intent is unclear or the user has just loaded the skill without describing their situation, open with the Mode Detection stage menu and ask where they are in the migration. Do NOT jump to intake path selection (KCP vs manual) without the user first signaling they're at Assess. Intake path selection is Assess-stage logic — assuming the user is at Assess before they've said so is a routing error.
  • Stage discipline. At each stage, address only that stage's decisions and immediate red flags. Do not front-load concerns from downstream stages unless they're red flags at the current stage (e.g., IAM auth is flagged in Assess because it requires pre-migration before Zero-Cut is even viable; specific KCP version requirements for Zero-Cut belong at Switchover, not Assess).
  • Command execution requires user approval — except for read-only operations on user-provided files. Mutating or environment-touching commands (KCP scans, Terraform, AWS CLI writes, Confluent CLI writes, anything that contacts an external system) require approval: present the command and ask whether the user wants to run it or have the skill run it. Do not auto-execute mutating commands without explicit approval. Read-only operations against files the user has explicitly pointed at — file reads, jq queries, structured parsing of state files / profiles / configs — auto-run with no approval prompt. Those are parsing, not executing, and asking permission to parse a file the user just provided breaks flow. The user is the migration practitioner and stays in control of their environment via the approval rule for mutating commands; read-only file parsing is not an environment-control concern.
  • One command per Bash tool call. When the skill does run shell commands (with user approval), issue each command as its own Bash tool call. Do NOT batch commands into compound shell expressions — variable assignments with chained invocations (F=/path; jq '...' "$F"; jq '...' "$F"), cmd1 && cmd2, subshells, loops over collections, heredocs. Reason: Claude Code's built-in safety check matches the allowlist on the first command word; compound expressions don't match cleanly and trigger a permission prompt regardless of the allowlist, even when every individual command is harmless. Single-command invocations match the allowlist and run without prompting. Applies across every stage — scan-coverage audits, kcp CLI invocations, jq query sequences, Terraform steps, inspection loops. Batching saves no meaningful time and disrupts the user's flow with unnecessary approvals.
  • Ask before acting on branching decisions. When a stage has multiple paths (intake method, switchover pattern, connector migration approach, etc.), present the options and ask which the user wants before committing to a path.
  • Foundational inputs — ask, never fabricate. Three inputs are load-bearing for the Plan's core recommendations: topics/partitions/scale, auth posture, and networking accessibility (private/public plus VPC topology). When one is missing, ask — route the user to a re-scan or manual intake — and hold the dependent sections rather than inventing a value. Throughput is foundational-but-degradable: no metrics at all → ask; peak present but P95 missing → use the existing peak fallback with the overestimation flag (do not hard-block). Everything else (EOS/transactions, Kafka Streams, connector detail, costs, client inventory, SR version, IBP, the finer reachability route) is peripheral — assume, label the working assumption, and capture an Open Question per the never-hedge behavior. The discriminator: if the assumption would invent a load-bearing recommendation (cluster type, sizing, networking, auth, switchover), ask; if it fills a peripheral unknown, assume and label. Do not trust a tool's success exit over the data: when a KCP state file shows kafka_admin_client_information.topics.details empty across all clusters while msk_cluster_config is populated, the deep scan did not complete (almost always private-network unreachability) — this is not a zero-topic cluster. Verify the state file actually contains topic data before proceeding.
  • Avoid temporal claims. No "as of Q1 2026," no release-date stamps, no "recent changes." Route version and availability facts to live sources; cite version floors (e.g., "v0.7.0+") without dates.
  • Full URLs in citation link text — strict. Every doc citation in skill output must render the full URL visibly in the rendered text, not just behind the markdown href. The link text [migrate-cc.md](https://...) fails this rule — readers who copy the Plan to plaintext, print it, paste it into Slack, or view it in a non-rendering tool see only the filename with no way to navigate to the doc. Three acceptable formats: (1) bare URL (preferred): Per Confluent docs, see https://docs.confluent.io/cloud/current/multi-cloud/cluster-linking/migrate-cc.md. Most markdown renderers auto-link bare URLs; plaintext readers still see the full path. (2) URL as link text: [https://docs.confluent.io/cloud/current/clusters/cluster-types.md](https://docs.confluent.io/cloud/current/clusters/cluster-types.md) — verbose but explicit. (3) Short descriptive label + bare URL in surrounding prose: Per the cluster types doc (https://docs.confluent.io/cloud/current/clusters/cluster-types.md). (These examples show the .md form as written in this skill; when you emit the citation to the user, swap the extension to .html per the fetch/cite rule in Sources of Truth.) Forbidden: filename-only or filename-shorthand link text where the href carries the full URL — [cluster-types.md](https://...), [migrate-cc.md](https://...), [aws-pni.md](https://...), See [private-networking.md](https://...). Also forbidden: bare bracketed shorthand with no link at all (e.g., ([cluster-types.html]) in table cells). Also forbidden: filename mentions in prose with no URL (e.g., "per private-networking.html"). The href being correct is not sufficient — the URL must be visible. Applies to all doc citations in Plan output, Assess output, and any other artifact the skill produces. Does not apply to KCP repo links where the repo name is the canonical identifier (e.g., confluentinc/kcp).

Mode Detection

This stage menu is the default opening when the user's intent is unclear. Present the three in-scope stages — Explore, Assess, Plan — let the user pick, and route from there. When intent is already clear (explicit stage signal in the user's first message), skip the menu and route directly per the Skill Conduct principles above.

When describing Assess in the opening, introduce KCP at the first mention of "scan." Beginners don't know what "scan" means in an MSK migration context — naming the tool grounds it. KCP is Confluent's open-source command-line tool for planning and executing Kafka migrations to Confluent Cloud (github.com/confluentinc/kcp). Acceptable opening phrasing for the Assess row: "Assess — scan your MSK environment with KCP (Confluent's open-source migration tool at github.com/confluentinc/kcp), or describe it manually if you don't have KCP installed yet. Surfaces red flags and builds an environment profile." Adjust phrasing for tone, but the load-bearing piece is that "scan" is paired with the tool that does the scanning. A bare "scan your MSK environment" without the KCP intro is not enough for a user who hasn't seen the skill before.

Explore is the lowest-commitment entry point. Many users open the skill with general questions before they're ready to scan or plan. Present Explore as a valid path — they don't have to start with Assess. Acceptable opening phrasing for the Explore row: "Explore — ask general questions about MSK and Confluent Cloud migration. Concepts (Cluster Linking, the KCP Gateway / Zero-Cut for no-restart cutovers, Schema Linking), feature comparisons, tooling, process. I'll cite sources from docs.confluent.io and the KCP repo." Gloss switchover-pattern jargon (Zero-Cut, Gateway, Dual-write, Manual CL) in plain English on first use, even in this menu — a first-time user does not know these terms. Lead with Cluster Linking, since plain Cluster Linking is the recommended default and the Gateway is the large/complex escalation.

User IntentStageRead
"I just have questions" / "what is X?" / "how does Y work?" / "explain Z" / "compare A vs B" / browsing-stage intentExploreThis file's "Explore stage" section below
"scan my MSK clusters" / "assess my environment" / "starting fresh"Assessreferences/assess.md
"what cluster type should I use" / "plan my migration"Planreferences/plan.md
"set up Cluster Linking" / "switch my clients over" / "monitor post-cutover" / "provision target cluster"Out of scope (Provision / Migrate / Switchover / Monitor)Decline and redirect to docs.confluent.io and the Confluent account team. Still offer Assess or Plan if useful.

If the user asks about KCP commands or MCP tools directly, use references/kcp-commands.md or references/mcp-integration.md respectively — reference material, not user-facing stages.

If overall intent is still unclear, start at Explore — let the user ask whatever they want, then route to Assess or Plan when they signal readiness. When the user signals an intent for downstream execution stages (Provision, Migrate, Switchover, Monitor) — for example "set up Cluster Linking", "switch my clients over", "monitor post-cutover" — acknowledge that those are out of scope for this skill and redirect to docs.confluent.io and the Confluent account team. Still complete an Assess or Plan if the user wants those, since Plan's switchover-approach and pre-migration-workstream decisions inform downstream execution.

Explore stage

Conversational Q&A about MSK-to-CC migration. No artifact produced. The user asks questions; the skill answers grounded in its references and live-fetched docs. Exit conditions: user signals readiness for Assess ("let's scan my environment", "I have a state file"), or signals readiness for Plan ("I have an environment profile already"), or ends the conversation.

What Explore covers (in scope):

  • Concepts: Cluster Linking (mirror-then-cutover, the recommended default), the KCP Gateway / Zero-Cut (a proxy that cuts clients over with no restart, for large/complex environments), Schema Linking, mTLS vs API Keys vs OAuth, PNI vs PrivateLink vs VPC peering vs Transit Gateway, eCKU vs CKU sizing, tiered storage, MSK Connect vs self-managed Connect, Debezium migration paths.
  • Comparisons: MSK Provisioned vs Confluent Cloud Enterprise/Dedicated, MSK Serverless vs Enterprise, IAM auth on MSK vs CC auth methods, MSK Glue Schema Registry vs CC Schema Registry.
  • Tooling: What KCP is and what it does, what the KCP commands produce, what CMU is, what the public cost estimator is for, what kcp create-asset migrate-acls iam does vs migrate-acls kafka.
  • Process: What a typical migration looks like, what stages exist, what pre-migration work is involved (IAM → SCRAM, schema migration order), what Zero-Cut prereqs are.
  • Specific feature questions: "Does CC Enterprise support mTLS on Azure?", "What's the Enterprise eCKU cap on PNI?", "What's the CL source Kafka version floor?" — answer by fetching the relevant docs.confluent.io page (using the .md URL pattern from the Fetch tool directive) and citing the value with the source.

What Explore does NOT cover:

  • Customer-specific recommendations without source data. "Should I use Enterprise or Dedicated for my cluster?" requires Assess (sizing math depends on throughput, partitions, ACLs from the source). Redirect: "That's an Assess question — I need your MSK environment details first. Want to scan with KCP or describe it manually?"
  • Non-MSK Kafka sources. The skill is MSK-only. Open-source Kafka / Confluent Platform / Aiven / Redpanda migrations are not in scope.
  • Greenfield Confluent Cloud setup. Migration only.
  • General Kafka programming questions unrelated to migration (producer/consumer code, Kafka Streams app development, schema design).
  • Pricing dollars. Per SKILL.md commercial-signals row — direct to the public cost estimator and the Confluent account team. Feature comparisons are fine; specific dollar figures are not.

Conduct in Explore:

  • Answer with citations. Every claim ties to a specific source — docs.confluent.io page (cite the URL), KCP repo file (cite the GitHub URL), this skill's reference files. Don't answer from training data alone — the migration product surface evolves; live sources are authoritative.
  • One question at a time. Answer what was asked. Don't pre-emptively dump the full reference file. If the user asks "what is Cluster Linking?", answer that — don't also explain Schema Linking and Zero-Cut unless they ask.
  • Offer the natural next step at the end of each answer. Examples: "Want me to walk through how Cluster Linking applies to your environment specifically? Tell me about your MSK setup and we can move to Assess." / "Curious about how this would look in your cutover? Once you've scanned with KCP, I can produce a Plan." Optional, low-pressure — the user may just want more questions answered.
  • Stay in scope. If the user asks about something outside MSK→CC migration (e.g., "how do I use Flink?"), acknowledge briefly and redirect — don't pull them out of the migration context unless they explicitly want to leave.
  • Honest about limits. "I don't have access to live customer data — Confluent's account team can confirm specifics for your org" is preferable to fabricating a customer-specific answer.

Routing out of Explore:

  • User signals scan readiness ("I have a state file at X" / "let's run a scan" / "ready to assess") → load references/assess.md and proceed.
  • User signals planning readiness ("I have an environment profile already" / "let's build a plan from this data") → load references/plan.md and proceed.
  • User asks an Assess-shaped question (cluster-type recommendation, sizing math, networking choice for their environment) → say so, offer to start Assess: "That depends on your environment. Want to scan with KCP or describe your setup manually?"

Intake Path Selection

Present both paths to the user and ask which they want. Do NOT auto-detect tool availability by running bash — per the command-execution principle, user approval is required before any command runs. Present the two options, let the user choose, and then proceed.

Path 1 — KCP deep scan. KCP is Confluent's open-source migration tool (github.com/confluentinc/kcp) for assessing AWS MSK environments and generating migration assets (Terraform, mirror-topic configs, migration orchestration) for Confluent Cloud. If you have it installed and AWS credentials for your MSK account, it can scan your environment and produce a kcp-state.json file that becomes the canonical profile for every downstream stage. Richest path when available.

Path 2 — Manual intake. For users who don't have KCP installed, don't have AWS credentials available, or are in a restricted AWS account / pre-engagement context. We walk through ~11 question groups and populate assets/migration-profile.yaml, which serves the same role downstream as a KCP state file — less detailed but workable.

If the user already has a KCP state file, treat that as completed KCP path output — skip directly to parsing it.

Once the user picks a path, load references/assess.md and follow that stage's reference for the next steps. Do not ask auxiliary questions or run any commands before the user has chosen.

First-mention reminder: whenever KCP is introduced in a conversation (including downstream stages), include the brief intro — who built it, that it's open source, the GitHub URL, what it does in the migration context. Don't assume the user knows.

Stage Workflow

Each stage has entry criteria, an exit artifact, and a reference file. Each stage validates its own work before handoff.

StageEntryExit Artifact (validation for this stage)Reference
ExploreUser has questions about MSK-to-CC migration concepts, comparisons, tooling, or processNo artifact — conversational Q&A grounded in skill references and live-fetched docs. Exit when user signals readiness for Assess or Plan, or ends the conversation.SKILL.md "Explore stage" section
AssessUser starts migrationEnvironment profile with required fields populated; red flags surfaced; when a KCP state file is provided with topics.details[] populated, a Topic-Level Readiness section classifies user topics into four buckets (Skip / Manual / Needs Config / Moves Cleanly) — see references/assess.md "Topic-Level Readiness" sectionreferences/assess.md
PlanEnvironment profile existsTechnical Plan output starts with the "About this Technical Plan" boilerplate from references/plan.md; architecture decisions documented; pre-migration requirements identifiedreferences/plan.md

Users may enter at any of the three stages. Explore is often the entry point for users who are still learning; Assess is the entry point for users who have an MSK environment to scan or describe; Plan is the entry point for users who already have an environment profile. The Plan exit artifact is the handoff for downstream execution stages (Provision, Migrate, Switchover, Monitor), which the user carries out with their Confluent account team using docs.confluent.io as the reference.

Cross-Stage Decision Logic

Cluster type: default to Enterprise, escalate to Dedicated only on hard limits

Enterprise is the recommended target for every migration. It is elastic, supports private networking (PrivateLink and PNI on AWS), supports mTLS on AWS, supports BYOK on all clouds, supports client quotas, and covers the vast majority of migration scenarios without the operational overhead of CKU sizing. Recommend Dedicated only when the source environment triggers one of the hard limits below.

How to use the hard-limits table:

  • ROUTE rows require a live fetch before answering. Do not answer an escalation question from memory. Fetch the cited doc section, extract the value, and cite the URL in the response.
  • HARDCODED row protocol. A HARDCODED row encodes a fact the skill could not route to a structured live source at design time. Treat the cached value as a working assumption, not a verified fact. Two sub-categories:
    • Doc-cited HARDCODED — a doc URL is given, but the fact lives in prose rather than a structured row. Before recommending, fetch the cited URL and check whether a structured representation now exists (table row, bullet list, version matrix). If yes, use the doc's value and tell the user "Skill's cached value was X; doc now says Y; using the doc value." If no structured representation exists, present the recommendation conditionally on the cached assumption — phrase it as "If [cached condition] still holds (the docs do not currently publish a structured matrix to confirm — see [URL]), then [recommendation]; otherwise [alternative]." Do NOT date-stamp the cached value to the user; the cited URL is the live-verification signal.
    • Uncited HARDCODED — no public doc URL exists for this fact (e.g., "no canonical public doc found"). Keep the "last verified YYYY-MM-DD" stamp on these rows. Without a URL the user can verify against, the date stamp is the only staleness signal they have. Still apply the conditional framing in the recommendation.
  • Multi-trigger profiles. When more than one row applies, apply each independently and cite every trigger. Do not stop at the first match.
  • Maintenance for uncited HARDCODED rows. Uncited rows have no public URL to drift-check against, so the date stamp is the only staleness signal. Review uncited HARDCODED rows on a quarterly cadence: confirm cached values with the relevant Confluent product team (Cluster Linking, Networking, Schema Registry, etc.) or internal product docs. Update both the cached value and the "last verified YYYY-MM-DD" stamp when reviewed. If a structured public doc emerges that exposes the fact, convert the row from HARDCODED to ROUTE.

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
58
Forks
11
Last commit
Sep 2026
Advanced
Item type
skill
Key
msk-migration
Source
github.com/confluentinc/agent-skills