Google Cloud VMs for nub

SkillCloud & infra

Lets your agent create, start, and use Google Cloud Linux or Windows virtual machines for work your local machine can't do.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Google Cloud VMs for nub skill

About this capability

Provision, start, reach, and use Google Cloud VMs for nub — for any real-OS work the local macOS host and Docker can't do (real Linux-kernel enforcement, real Windows/AppContainer/MSVC, a clean multi-GB build box). Invoke whenever you think "I need a Linux box" or "I need a Windows box" — you can ST

What this skill tells your AI

The instructions your AI receives, as published by nubjs/nub in .claude/skills/gcloud-vm/SKILL.md and read by ahel’s review.

Any time a task needs a real OS the Mac host + Docker can't give you — real Linux-kernel Landlock/seccomp/netns enforcement, real Windows AppContainer/MSVC, a clean high-RAM build box, a genuinely-clean first-run environment — spin up or start a VM. Project pullfrog; the existing boxes live in us-central1-a. The gcloud default zone is us-west1-a, so always pass --zone us-central1-a explicitly.

⛔⛔ THESE ARE GOOGLE CLOUD BOXES — THE AWS FREEZE DOES NOT TOUCH THEM

Project pullfrog. An instruction that "the AWS account is frozen" says NOTHING about these, and reading it as covering them cost a whole session of 20-minute CI round-trips while two admin-capable Windows VMs sat idle and billing. The VMs are almost never the blocker — check before you route around them.

The standing instances (gcloud compute instances list is the truth; this table rots)

NameOSPurpose
nub-linuxUbuntu 24.04 LTS, e2-standard-4Linux-kernel enforcement (Landlock/seccomp/bwrap/netns); carries a path-bound AppArmor bwrap-userns profile + apparmor_restrict_unprivileged_userns=1 reproducing a locked-down 24.04 host. The one box worth preserving — that profile is hand-built, so it auto-STOPs rather than auto-deletes. It carries an 8h maxRunDuration from before that flag was dropped for new boxes; that cap CANNOT be raised while it runs, so a job here longer than 8h gets stopped mid-flight. STOP keeps the disk, so restart and resume rather than losing it

nub-corpus-linux, nub-win2 and nub-win3 were DELETED 2026-09-01 as part of the billing cleanup below, along with nub-win, nub-devloop and nub-win4 before them. Do not go looking for them, and do not treat their absence as an outage: recreate what you need from the recipes below, at the spec the job actually needs. The Windows boxes in particular were confirmed admin (IsInRole(Administrator)=True) with logman/wpr/tracerpt present, so a fresh Server 2022 box gives you full ETW kernel tracing the same way.

They are usually TERMINATED to save billing. Start what you need, and DELETE it when done — deleting is what stops the disk charge.

⛔ DO NOT PUT --max-run-duration ON A BOX YOU CREATE (reversed 2026-09-04)

A run-duration budget cannot be changed while the instance is running, and that is disqualifying. Measured 2026-09-04, on a box holding a 1h25m acceptance sweep that needed more time than its budget allowed:

ERROR: Max run duration cannot be changed while the instance is running.

gcloud compute instances update does not accept the flag at all; set-scheduling accepts it and the API refuses. --termination-time is refused the same way, and clearing the duration counts as changing it. The only way out is to stop the instance — which kills the work the box exists for. So the flag turns a recoverable "this is taking longer than I thought" into an unrecoverable one, and it does it at the exact moment the box is most valuable.

Use idle auto-stop instead (below). It targets the actual waste — a box sitting doing nothing — rather than capping useful work, and it can be tuned from inside the guest at any time.

The cost facts that motivated the budget are still true, and still worth knowing. Measured 2026-09-01: five boxes alive, four idle between 10 hours and 26 days, carrying $94/mo of disk that bills while TERMINATED. nub-win3 alone was ~$205/mo of runtime — 322 running hours in 30 days, of which $119 was Windows licensing rather than compute — on top of $34/mo for its 200 GB pd-ssd.

  • DELETE is what stops the standing charge. The boot disk is autoDelete by default, so deleting the VM takes it too; merely stopping one leaves a 200 GB pd-ssd billing ~$34/mo forever. Delete anything named -tmp, -probe or a builder the moment it is done.
  • Windows is where the money is. The licence is $0.046 per vCPU-hour, so an e2-standard-8 Windows box costs more in licence than in compute ($0.368/h vs $0.268/h). Size Windows boxes by what the job needs, and prefer --boot-disk-type pd-balanced ($0.10/GB/mo) over pd-ssd ($0.17) unless the disk is genuinely the bottleneck.
  • Verify a cleanup by listing EVERY instance and reading the list — never by grepping for the name you expected. A cleanup check that matched nub-win passed while its own box, named wingrants-…, ran on for hours.

Automated short-lived builders under scripts/remote-build.ts are a separate case and still carry the flag: they run ~45m unattended, nobody is there to extend anything, and the launcher dying is a real leak path.

Idle auto-stop — the mechanism that replaces a run-duration cap

This is the cost mechanism, now that the run-duration cap is gone: it targets idleness directly rather than capping useful work. Ship it in the same create call so no box exists without it, and the box powers itself off once nothing is happening. GCE reports a guest poweroff as TERMINATED, so the disk survives and a start brings it back.

cat > /tmp/idle-stop.sh <<'SH'
#!/usr/bin/env bash
# power off once nothing has happened for $IDLE_MIN minutes
set -euo pipefail
IDLE_MIN=30
STAMP=/var/tmp/nub-idle-since
busy() {
  who | grep -q .                     && return 0   # an SSH session is attached
  pgrep -x 'cargo|rustc|nub' >/dev/null && return 0  # a build is running detached
  awk '{exit ($1 > 0.2) ? 0 : 1}' /proc/loadavg      # 1-min load says real work
}
if busy; then rm -f "$STAMP"; exit 0; fi
now=$(date +%s); [ -f "$STAMP" ] || printf '%s' "$now" > "$STAMP"
if [ $(( (now - $(cat "$STAMP")) / 60 )) -ge "$IDLE_MIN" ]; then
  logger -t nub-idle "idle ${IDLE_MIN}m, powering off"
  systemctl poweroff
fi
SH

cat > /tmp/idle-startup.sh <<'SH'
#!/bin/bash
install -m 0755 /dev/stdin /usr/local/sbin/nub-idle-stop <<'INNER'
__IDLE_SCRIPT__
INNER
cat > /etc/systemd/system/nub-idle.service <<'UNIT'
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/nub-idle-stop
UNIT
cat > /etc/systemd/system/nub-idle.timer <<'UNIT'
[Timer]
OnBootSec=10min
OnUnitActiveSec=5min
[Install]
WantedBy=timers.target
UNIT
systemctl daemon-reload && systemctl enable --now nub-idle.timer
SH
python3 - <<'PY'
p='/tmp/idle-startup.sh'; s=open(p).read()
open(p,'w').write(s.replace('__IDLE_SCRIPT__', open('/tmp/idle-stop.sh').read().rstrip()))
PY
# then add to the create call:
#   --metadata-from-file startup-script=/tmp/idle-startup.sh
  • who is the load-bearing check, because an agent driving a long cargo build over SSH holds a session the whole time — the timer must not shoot the box out from under it. The pgrep arm covers a build detached from any session, and the load arm covers everything else.
  • Tune IDLE_MIN up, never down, for a box you interact with by hand. Thirty minutes is chosen so a poweroff never lands mid-thought. There is no run-duration backstop behind it any more, so a wedged timer means a box that bills until someone deletes it — check instances list when you finish with a box.
  • Windows has no equivalent here. Nothing reaps a Windows box for you, so DELETE it explicitly when the job is done and confirm with a full instances list.

⛔ LIFECYCLE HYGIENE — DIAGNOSE AT FIRST DETECTION, THEN DELETE

These boxes are THROWAWAY. The maintainer's standing instruction: kill anything unreachable rather than nursing it, and create a fresh one at whatever spec the job needs. A borked box that keeps running is pure burn.

The rule that actually matters: the moment you find a box unreachable, DIAGNOSE IT THEN — not later. Once it is deleted, or once weeks pass, every trace of what went wrong is gone and you are left guessing. Read the serial console (gcloud compute instances get-serial-port-output <name> --zone us-central1-a | tail -40) BEFORE deleting, and write down what you find.

gcloud compute instances delete <name> --zone us-central1-a --quiet   # a stopped box still bills its disk

⛔ A FULL DISK IS INDISTINGUISHABLE FROM A BROKEN BOX, AND IT IS THE MOST COMMON CAUSE HERE. Measured 2026-08-05 on nub-linux: /dev/root 193G 193G 0 100%, of which 166 GB was abandoned ~/.cache/nub-search-* harness fixture roots (8 of them) — no runaway process, just temp dirs nothing ever swept. The symptoms all look like a dead machine: scp: write remote "x": Failure, zero-byte outputs from commands that "succeeded", a transferred file that reads as cannot execute binary file. Check df -h / FIRST on any box behaving strangely; rm -rf ~/.cache/nub-search-* took it from 100% to 14% and fully restored the box, no recreation needed.

Re-test reachability rather than trusting a remembered "unreachable". The external IP changes on every start, so a stale IP reads exactly like a dead box.

gcloud compute instances list                                   # names + STATUS + current external IP
gcloud compute instances start nub-linux --zone us-central1-a   # ~30-60s; Windows boot is slower
gcloud compute instances stop  nub-linux --zone us-central1-a   # when you finish — they bill while RUNNING

SSH — user nub, key ~/.ssh/nub-vm, and the IP is DYNAMIC

  • User is nub — not ubuntu/nubuser/colinmcd94 (all fail Permission denied (publickey)).
  • Key is ~/.ssh/nub-vm.
  • The external IP changes on every start (no static IP reserved). Never hardcode it or trust a memory/skill note:
    IP=$(gcloud compute instances describe nub-win --zone us-central1-a \
          --format='value(networkInterfaces[0].accessConfigs[0].natIP)')
    ssh -i ~/.ssh/nub-vm -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null \
        -o ConnectTimeout=15 nub@"$IP" "echo ok"
    
  • Re-resolve the IP on every reconnect, not once per session. A VM can be stopped out from under you mid-session and return on a different IP, so an address resolved at the top of a long run can be dead an hour later. Treat a sudden Connection refused/timeout on a previously-working box as "it moved" and re-read the IP before reading the serial console.
  • Always reachability-guard a VM dispatch (ConnectTimeout, timeout), and after a fresh start retry with backoff for a few minutes — sshd (Windows especially) isn't up the instant STATUS flips RUNNING. Never let a sub-agent hang on a VM: check reachability, act, report, exit.

Creating a NEW VM on demand

For an isolated/ephemeral box (a clean first-run env, a second Linux box so you don't contend with nub-linux, a specific image), create one — don't wait for permission:

# Linux — size ≥16 GB if it will COMPILE nub (see the OOM gotcha); e2-standard-4 is the proven size.
# The last two lines are NOT optional: they are what stop a forgotten box billing for weeks.
gcloud compute instances create nub-linux-tmp \
  --zone us-central1-a --project pullfrog \
  --machine-type e2-standard-4 \
  --image-family ubuntu-2404-lts-amd64 --image-project ubuntu-os-cloud \
  --boot-disk-size 30GB \
  --metadata-from-file startup-script=/tmp/idle-startup.sh

# Wire the `nub` SSH key so you can reach it the same way (Linux):
gcloud compute instances add-metadata nub-linux-tmp --zone us-central1-a \
  --metadata ssh-keys="nub:$(cat ~/.ssh/nub-vm.pub)"

Windows: ssh-keys metadata alone does NOT get you in — provision it yourself

Measured end-to-end 2026-08-04 while building nub-win3, after nub-win and nub-win2 both proved unreachable. A Windows box created the "obvious" way is not SSH-able, and each of the three failures below looks like a different problem, so use the failure MODE to tell them apart rather than guessing:

symptommeaning
Operation timed outOpenSSH Server is not installed — Windows Server 2022 does not ship it enabled, and enable-windows-ssh=TRUE does not install it
Connection refusedsshd installed but not started yet (still booting)
Permission denied (publickey…)sshd is up and listening; only the key is missing

Do it in ONE creation, with a startup script that installs sshd and provisions the key itself. The guest agent did not provision keys here across two resets, so do not depend on it:

cat > /tmp/win-ssh.ps1 <<PS
Add-WindowsCapability -Online -Name OpenSSH.Server~~~~0.0.1.0
Set-Service -Name sshd -StartupType Automatic
Start-Service sshd
\$pw = ConvertTo-SecureString (([guid]::NewGuid()).ToString() + '!Aa1') -AsPlainText -Force
if (-not (Get-LocalUser -Name nub -ErrorAction SilentlyContinue)) {
  New-LocalUser -Name nub -Password \$pw -PasswordNeverExpires -AccountNeverExpires
}
Add-LocalGroupMember -Group Administrators -Member nub -ErrorAction SilentlyContinue
# ⛔ An ADMIN user authenticates via administrators_authorized_keys — a key in the user's own
# ~/.ssh/authorized_keys is SILENTLY IGNORED, which is what "Permission denied" was really saying.
\$ak = 'C:\ProgramData\ssh\administrators_authorized_keys'
Set-Content -Path \$ak -Value '$(cat ~/.ssh/nub-vm.pub)' -Encoding ascii
icacls \$ak /inheritance:r
icacls \$ak /grant 'Administrators:F' /grant 'SYSTEM:F'   # sshd REFUSES a loosely-ACL'd key file
Restart-Service sshd
PS

# There is no guest-side idle timer for Windows and no run-duration cap, and the licence makes this
# the most expensive box shape in the project — so delete it explicitly the moment the job is done.
gcloud compute instances create nub-win-tmp \
  --zone us-central1-a --project pullfrog \
  --machine-type e2-standard-8 \
  --image-family windows-2022 --image-project windows-cloud \
  --boot-disk-size 200GB --boot-disk-type pd-ssd \
  --metadata-from-file windows-startup-script-ps1=/tmp/win-ssh.ps1
  • Budget ~10-15 minutes from create to first successful SSH; first boot plus sysprep is slow, and the startup script runs partway through it. Poll rather than waiting on one attempt.
  • The default SSH shell is cmd.exe, not PowerShell. A ;-separated PowerShell one-liner dies with Invalid argument/option - ';'. Wrap it: ssh … 'powershell -NoProfile -Command "…"'.
  • Disk: 200 GB. The image is 50 GB and gcloud warns the root partition may need manual resizing; Server 2022 resized it automatically here (179 GB free on first login). A debug box that builds nub and installs npm trees fills a 50 GB disk.
  • Read the serial console to confirm the script ranget-serial-port-output … | grep windows-startup-script-ps1 shows each line's output, including processed file: C:\ProgramData\ssh\administrators_authorized_keys.
  • A firewall rule is NOT the problem: default-allow-ssh is 0.0.0.0/0 tcp:22 with no target tags, so it already covers every instance. Don't go hunting network tags.

Running a LONG job (a cargo build) on a Windows box — use a scheduled task, and scp the script

A cold cargo build --release -p nub-cli is ~30-45 min, far past any single SSH call. Four approaches were tried on nub-win3; three failed, and the failure MODES are the useful part because two of them are SILENT:

approachwhat happened
Start-Process … -WindowStyle Hiddenran, then died after the dependency downloads with NO error line. The SSH session's job object closes and takes the child with it. A vanished process with an empty log is this, not a build error.
multi-line PowerShell piped to powershell -Command - over SSH stdinthe script silently never materialised (if exist … NO_BAT). Quoting dies somewhere between zsh, ssh and PowerShell.
scheduled task running as SYSTEM, cargo via ~/.cargo/bin/cargo.exeerror: rustup could not choose a version of cargo to runEXIT=1. rustup's default toolchain is per-USER, and SYSTEM has none.
scp the .bat, then run it as a scheduled task, calling the toolchain binary directlyworks
# Write the .bat LOCALLY with CRLF and scp it — do NOT try to author it over SSH stdin.
printf '@echo off\r\ncd /d C:\\nub\r\nset RUSTUP_HOME=C:\\Users\\nub\\.rustup\r\nset CARGO_HOME=C:\\Users\\nub\\.cargo\r\n"C:\\Users\\nub\\.rustup\\toolchains\\stable-x86_64-pc-windows-msvc\\bin\\cargo.exe" build --release -p nub-cli > C:\\nub\\build.log 2>&1\r\necho EXIT=%%ERRORLEVEL%% >> C:\\nub\\build.log\r\n' > /tmp/dobuild.bat
scp -i ~/.ssh/nub-vm /tmp/dobuild.bat nub@"$IP":C:/nub/dobuild.bat
# /RU SYSTEM is fine HERE because a build only needs a toolchain — never reuse this line for a measurement.
ssh -i ~/.ssh/nub-vm nub@"$IP" 'cmd /c "schtasks /Create /TN nubbuild /TR C:\nub\dobuild.bat /SC ONCE /ST 00:00 /RL HIGHEST /RU SYSTEM /F & schtasks /Run /TN nubbuild"'
# then POLL: (Get-Process cargo,rustc).Count, plus the tail of build.log

⛔⛔ USE THIS FOR BUILDS ONLY — NEVER FOR A MEASUREMENT. A scheduled task runs as SYSTEM, and SYSTEM IS NOT A NORMAL USER. For a build that costs a PATH fix (the rustup row above). For anything that MEASURES OS-enforced behaviour it silently changes the answer, because SYSTEM holds privileges an ordinary account does not — SeCreateSymbolicLinkPrivilege above all — and os.homedir() becomes C:\Windows\system32\config\systemprofile.

Measured 2026-08-04, and the verdict did not give it away: a build-jail package whose real failure is a refused symlink was re-measured under a SYSTEM task, produced exactly the expected grant with a clean control, and was reported as a validated reproduction. The artifact refuted it — the per-cell log was 1,476 bytes containing only a catalog warning naming the systemprofile path, and the symlink error appeared in 3 of 54 logs where the real-user CI run shows 51 of 54. Running as SYSTEM had bypassed the very mechanism under test while landing on the same answer by another route.

  • whoami is the cheap guard. Print it as the first line of any script whose result you will believe, and assert on it: nub-win3\nub good, nt authority\system void.
  • Get the right context by running through SSH itself, not a task — an SSH session already runs as nub. Wrap the call in a harness-tracked background command (run_in_background) rather than a scheduled task: the session stays alive for hours, so the ~10-minute foreground cap that pushed you toward schtasks never applies. Detaching within the session (Start-Process, nohup-alikes) still dies to the job object — the point is that the SSH call itself is the long-lived process.
  • schtasks /RU <user> /RP <password> also works but needs a password you probably do not have — the nub account is created with a throwaway GUID password, and net user nub <new> fails The user name or password is incorrect from a non-elevated SSH token. Prefer the SSH route.
  • Call the TOOLCHAIN binary, not the rustup shim (.rustup\toolchains\stable-x86_64-pc-windows-msvc\bin\cargo.exe) so it does not matter which user rustup was configured for.
  • The scheduled task is what makes failure VISIBLE — it redirects to a log and records EXIT=<n>, where the detached-process approach just disappears.
  • Toolchain prerequisites, ~15 min before any build: VS Build Tools (--add Microsoft.VisualStudio.Workload.VCTools --add Microsoft.VisualStudio.Component.Windows11SDK.22621 --includeRecommended) then rustup-init.exe -y --default-toolchain stable --profile minimal. Verify by running cargo --version from its absolute path, not by trusting the installer's exit.
  • Node/git come from vendor installers, silently: node .msi via msiexec /qn, Git-for-Windows .exe via /VERYSILENT /NORESTART. Neither is on PATH for an existing SSH session — read [Environment]::GetEnvironmentVariable("Path","Machine") or use absolute paths (C:\Program Files\Git\cmd\git.exe).

Delete an ephemeral box when done — a created VM keeps billing its disk even when stopped:

gcloud compute instances delete <name> --zone us-central1-a --quiet

Auth — the service-account key is the durable path

The USER credential (colin@pullfrog.com) has its refresh token revoked periodically by org session-control policy, so gcloud auth login is not durable. A service-account key is exempt and works non-interactively without changing gcloud's global state:

CLOUDSDK_AUTH_CREDENTIAL_FILE_OVERRIDE=~/.config/pullfrog/vertex-service-account.json \
  gcloud compute instances list --project=pullfrog

The SA (pullfrog-vertex-e2e@pullfrog.iam.gserviceaccount.com, project pullfrog) has Owner, so list/describe/start/stop/create all work through the override. This is the preferred path. Fall back to ! gcloud auth login only if the override itself errors Reauthentication failed. cannot prompt during non-interactive execution (key removed/rotated).

That Owner grant is wider than this skill needs, and it is shared with something else — do not "fix" it by narrowing this SA. The same key is the VERTEX_SERVICE_ACCOUNT_JSON GitHub Actions secret in pullfrog/app, where it is used only to authenticate Gemini inference. Narrowing it to roles/aiplatform.user — the obvious tightening if you look at the pullfrog side alone — silently breaks every VM operation in this skill, because that role carries no compute permission at all. The correct shape is two accounts: leave CI with a Vertex-only SA, and mint a separate roles/compute.instanceAdmin.v1 key for VM ops that never enters CI. Until that split exists, treat this key as an operator credential and keep it off any surface an untrusted run can read.

Gotchas

  • A RUNNING instance can be a DEAD instance — read the serial console FIRST:
    gcloud compute instances get-serial-port-output nub-linux --zone us-central1-a | tail -40
    
    A wedged box typically shows Out of memory: Killed process (rustc). Serial output beats guessing "network problem."
  • Size ≥16 GB for anything that compiles the nub Rust workspace. An e2-small (2 GB) cannot build it and will OOM-wedge. e2-standard-4 (16 GB) is the proven size.
  • Write every script you send to nub-win as ASCII + CRLF. PowerShell 5.1 reads a BOM-less script in the ANSI codepage, so a UTF-8 character anywhere in the file — an em-dash in a comment is the usual culprit — fails with The string is missing the terminator, and the error points at the last line of the file, not the offending one. Keep remote PowerShell ASCII-only, or emit a UTF-8 BOM.
  • IsOutputRedirected is always True over SSH, so the first-run TTY path is unreachable there. Any is_terminal() branch silently takes the non-TTY path. Testing real console behavior on Windows needs a ConPTY harness or an interactive RDP session. (On Linux, wrap the run in script(1) and the TTY branch runs.)
  • For a nub RUST BUILD, use the remote-build skill, not this one. scripts/remote-build.ts provisions an ephemeral spot builder from a pre-baked image and cross-compiles aarch64-apple-darwin on Linux. This skill remains the entry point for Windows/MSVC and for an interactive box.
  • Prefer cross-compile-on-Mac + scp the artifact over building on the VM. Running a binary needs almost no RAM. For Windows, the VM's MSVC BuildTools is often a broken shell with no cl.exe — cross-compile for x86_64-pc-windows-gnu on the Mac (rustup target add …; brew install mingw-w64 if the linker is missing), strip, scp the .exe, run it. harness = false test binaries are self-contained and ideal for this. The Windows home dir may be C:/Users/nub.<HOST>/, not C:/Users/nub.
  • Judge results by behavioral/differential evidence (EPERM vs success, byte counts, a before/after delta), not wall-clock — a shared VM may be contended.

Related

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
4k
Forks
60
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
gcloud-vm
Source
github.com/nubjs/nub