Submitting with the VCC CLI
The vcc command-line tool lets you download Challenge data, validate and package your predictions, and submit results to the Virtual Cell Challenge directly from your terminal or remote environment — no browser upload required. It works on macOS, Linux, and Windows, and in headless or remote environments where a browser upload isn't practical.
📦 On PyPI: vcc-cli — uv tool install vcc-cli installs the vcc command.
Before you start
You'll need a Challenge account and an API key. Sign in at the Challenge portal, complete registration, and generate your key from Account → Credentials. Keys are shown once — if you lose it, just generate a new one (this revokes the old key; one active key per account).
1. Install
Prerequisites
The CLI is installed with uv. Skip anything you already have.
macOS
/bin/bash -c "$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)" # Homebrew, skip if `brew` works
brew install uvLinux (no sudo needed)
curl -LsSf https://astral.sh/uv/install.sh | sh # installs to ~/.localWindows (PowerShell)
powershell -c "irm https://astral.sh/uv/install.ps1 | iex"Install the CLI
uv tool install vcc-cli # the command is still `vcc`
vcc --versionvcc --version should print the CLI version and your Python version.
command not found: vcc?
It's installed but not on your PATH. Run uv tool update-shell, open a new terminal, and retry. Don't reinstall.
To update later: uv tool upgrade vcc-cli.
Optional: install the agent skill
If you use a coding agent (Claude Code, Codex, Gemini), the CLI can install a skill so your agent can drive the whole workflow for you:
vcc skill installThen restart your agent session (or run /reload). After that you can simply ask your agent, e.g. "download the VCC control data and make a submission." Re-run vcc skill install after every CLI upgrade so the skill matches the installed version.
2. Authenticate
Generate your API key in the member portal (Account → Credentials), then log in immediately so the one-time key isn't lost.
macOS / Windows
vcc login --token-stdinPaste the key at the hidden prompt. The output should say stored: OS keychain. If a keychain prompt appears, click Always Allow.
Linux
vcc login --token-stdinIf your machine has a keyring, this works the same as above. On headless or shared machines (common on clusters), vcc login will refuse to store the key insecurely — that's intentional. Instead, export it as an environment variable and skip vcc login entirely:
export VCC_TOKEN=<your-api-key>The exported token authenticates every command.
Verify
vcc whoamiThis shows your account, team, endpoint, and a masked token, plus either ready to submit or NOT able to submit yet with the specific blockers (e.g. registration not yet approved).
3. Download the control data
vcc datasets list
vcc datasets download controls -d ~/vccThe controls bundle is several hundred MB; vcc datasets list shows the exact size before you start. The download shows a live rate and ETA and ends with a checksum verification (verified: crc32c checksum matches). Downloads are resumable — if it's interrupted, re-run the same command and it picks up where it left off.
The 2026 Challenge has three cell contexts
The controls bundle contains one control .h5ad per context — context_A.h5ad, context_B.h5ad, context_C.h5ad — plus the shared gene_names.csv and pert_counts.csv. Your model predicts the same perturbations in all three contexts, and you submit one .vcc file covering all of them, with each cell tagged A, B, or C in a context column. Reuse those labels exactly as they appear in the control files: they are how your predictions are matched to the right held-out data at scoring time.
What a "context" is, and how you tell them apart
Each context is a different cell line. We don't publish which one. A, B and C are deliberately opaque labels.
What you get instead is better than a name: 18,400 non-targeting control cells per context, in that context's control file. That unperturbed expression profile is the context — it's what tells your model which cellular state it's predicting a perturbation in. Two contexts might respond differently to the same knockout, which is the point of the Challenge; the control cells are how you tell which is which.
Never reorder or reassign the context labels
The label on a cell decides which held-out dataset it's compared against. If you swap two contexts — by relabelling, by concatenating files in a different order than you think, or by generating your own controls — every metric degrades toward chance and your result looks like a weak model rather than a bug. That is the most expensive mistake available here, because nothing about the score says "you mixed up the labels".
Keep the label attached to the cells from the moment you read a control file. vcc prep checks that all three contexts are present and that none is unknown, but it cannot tell that A and B were swapped — only you can prevent that.
The final phase uses different labels: D, E, F
The final test set is three different cell lines, released as contexts D, E and F with their own control files. The labels don't overlap with A/B/C on purpose: if they did, a final submission that accidentally reused validation labels would be silently scored against the wrong cell lines. Use whichever labels came in the bundle you're predicting for, and don't carry A/B/C into a final phase submission.
The two rounds are scored separately and are not comparable — different cell lines, different perturbation panels, and one of the metrics is a rank within its own panel. A lower final phase score than your validation score doesn't mean your model got worse.
All files sit at the root of the zip — there is no vcc_data/ directory:
unzip -o -j ~/vcc/vcc_2026_controls.zip "gene_names.csv" "pert_counts.csv" "manifest.json" -d ~/vcc4. Package your prediction
Predictions are submitted as .vcc files. Convert your AnnData prediction with vcc prep, which needs both reference files from the controls bundle — the gene list and the perturbation list:
vcc prep ~/path/to/prediction.h5ad \
-g ~/vcc/gene_names.csv \
--perts ~/vcc/pert_counts.csv \
-o ~/vcc/prediction.vccprep validates your file against the Challenge format — gene set, contexts, perturbation labels, per-perturbation cell counts, raw-counts check, and the no-control-cells check — before packaging, so format problems surface locally in seconds instead of after a multi-GB upload. Both gene_names.csv and pert_counts.csv are included in the controls bundle.
The perturbation list is --perts, with no short form
Don't copy -p across from vcc sample. In prep, -p is --pert-col, a column name — see the -p gotcha.
Add --dry-run to validate without writing a file. On success prep reports the cells per context and confirms targets: verified against the official list.
Just want to test the pipeline?
vcc sample generates a random-data submission so you can exercise the whole pipeline without a real model. It scores poorly by design.
vcc sample -g ~/vcc/gene_names.csv -p ~/vcc/pert_counts.csv -o sample.vccThe -p flag is required — without it the file contains placeholder perturbations and the scorer will reject it with a "perturbation mismatch" error. With it, the sample gives every perturbation the official 400 cells — the count is built into the CLI, since pert_counts.csv is a bare list of perturbations — so it satisfies the cell-count check and the scorer accepts it.
5. Submit
vcc submit ~/vcc/prediction.vcc -m "my model v3" --waitThe flow: submission-limit check → entry created (note the entry id) → upload with progress → scoring started. With --wait, the CLI polls through launching → scoring → published and prints your rank, overall score, and the six component metrics:
Rank: 2
Overall: 0.3614
Perturbation discrimination pds: 0.6222
Expression accuracy mse: 0.0
DE log-FC accuracy nmae: 0.4137
DE direction fidelity fid: 0.3312
DE direction reach reach: 0.2911
DE significance overlap jac: 0.5102
Partition: val
Panel: vcc2026-val-1
Anchor set: vcc2026-valA-r1+vcc2026-valB-r1+vcc2026-valC-r1The short name beside each metric (pds, mse, …) is the column header for that metric on the leaderboard, so the score you read here and the column you find your rank in are named the same way. The six average to Overall exactly — that is what "unweighted mean" means, and you can check it with a calculator.
How to read these. Each of the six is on a common scale with two named points: 0 is the cell-context-mean baseline — no better than pasting the average cell onto every perturbation — and 1 is a split-half replicate of the real experiment, i.e. as good as one half of the real data predicts the other. Overall is their mean.
⚠️ 1.0 is a landmark, not a maximum. A perfect prediction beats a split-half replicate on most of the metrics, because the replicate carries the real experiment's own noise. Only Expression accuracy is capped at 1.0; the other five are unbounded above. So a score above 1 is possible and means "better than a replicate", and the scale is not a percentage of anything — don't read 0.36 as "36% correct".
There is no fixed floor, so a negative score is a real value too: it means worse than the baseline, which is what a random or mis-labelled submission looks like. How far below 0 a metric can go differs by metric — Expression accuracy stops at 0, DE log-FC accuracy is floored at −6 by a penalty cap, and the other four bottom out at their own depths (shallow for DE significance overlap, around −0.1; deeper for DE direction fidelity, about −1.9).
The last three lines are the stamp: which partition, perturbation panel and scoring bundles produced the numbers. Anchor set names one bundle per cell context — each context is scaled against its own two reference points — so it is an identity string, not a version number. Scores are not comparable across any of the three, so a score is not readable without them.
Raw cell-eval2 values (pds_cosine, de_wilcoxon_lfc_nmae, …) are deliberately not printed — they sit on six different scales and two point the other way, so showing them in the same column invites reading a raw 5.25 as a score. They are still in vcc status --json.
Submissions count against your daily limit, renewing at midnight UTC. Only submissions that reach scoring count: usage errors, a 409 refusal, and anything rejected before a score is saved are free. Your team can have only one submission in flight at a time; a second one returns HTTP 409 until the first finishes.
Interrupted upload?
Re-run with --resume to continue the same entry instead of creating a new one:
vcc submit ~/vcc/prediction.vcc --resume6. Check status any time
vcc status <entry-id>Shows the entry's current status and scores once published. Add --wait to stream until it reaches a terminal state. A failed entry exits non-zero, which makes it easy to use in scripts.
Logging out
You normally never need to log out. If you're on a shared machine and want to remove the stored credential:
vcc logoutSubmission requirements (2026)
A submission is a single .vcc file covering all three cell contexts. vcc prep checks every requirement below before anything is uploaded, so problems surface on your machine in seconds rather than after a multi-GB transfer. The same rules are enforced when your file is scored, so a submission that passes prep will not be rejected later for format reasons.
| Requirement | Detail |
|---|---|
| All three contexts, one file | Every cell carries a context label — reusing the labels from the control files you downloaded (A/B/C during the validation phase, D/E/F in the final phase). One file contains all three; do not submit per-context files. |
| Exact perturbation set | Each context predicts exactly the perturbations listed in pert_counts.csv — one column, one row per perturbation, the same set in every context; no extras, none missing. Labels are gene symbols (ADNP), not construct ids (ADNP-1). An extra perturbation is rejected, not ignored: one of the metrics is computed over the whole scored set, so a submission with extras isn't being scored on the same thing as everyone else. |
| Exact cell counts | Every perturbation has exactly 400 cells, in every context — the same number for all of them. Not a minimum: one cell over or under is rejected. It is not the same as the 2025 Challenge's count, and vcc prep checks it for you (--cells-per-pert overrides it if a future bundle ever changes). |
| Exact gene set | All 18,533 genes from gene_names.csv, no more and no fewer. Order is fixed up for you automatically. |
| Raw counts | Submit the raw count matrix, not log-normalized data. Every value must be a whole number and non-negative (and finite) — scoring runs in counts space and rejects a fractional matrix outright. Whole numbers stored as floats are fine; you do not need an integer dtype. If your model emits continuous values, round them before prepping. |
| Per-cell count cap | No single cell may total more than 1,000,000 counts (summed across all genes). vcc prep checks this (as does scoring), so an over-cap cell fails locally rather than after upload. --max-counts-per-cell -1 disables the local check. |
| Total cell cap | A submission may not exceed 400,000 cells in total. The 2026 panel comes to 360,000, so this only bites if you have added cells that don't belong. |
| Density cap | At most 4,750,000,000 stored entries across the whole prediction — about 13,200 per cell at the 360,000-cell panel. A separate ceiling from the cell cap, and not a file-size limit: predictions compress well, so a 25 MB .vcc can need more memory than a 3.3 GB one. It counts what the matrix stores, not what is mathematically non-zero, so a zero you saved explicitly still counts — which is why a dense array (6,671,880,000 entries, 1.40× the cap) is over the limit on its own. |
| No control cells | Do not include non-targeting cells — a submission containing them is rejected, not cleaned up for you. Scoring pairs each context with the held-out control cells rather than yours, which is why you can omit them entirely. |
A complete submission is therefore 300 perturbations × 400 cells × 3 contexts = 360,000 cells, over 18,533 genes.
Too dense? Predict fewer expressed genes — don't drop cells
The densest thing anyone has a real reason to submit is the baseline: predict the context mean for every perturbation, which leaves nearly two thirds of the genes nonzero. That comes to about 11,800 nonzeros per cell — inside the cap, but with only about 10% to spare. A matrix stored densely holds all 18,533 genes for every cell, roughly 40% over the cap by itself, so that is nearly always the cause — saving as scipy.sparse.csr_matrix before vcc prep fixes it outright.
If you are genuinely near the limit, zero out genes below ~1 CPM — counts per million, per cell. That costs 0.07% of the expression mass, and the DE metrics already discard everything under 5 CPM. Do not drop cells instead — the per-perturbation cell counts are exact, and a short submission is rejected.
--max-nnz -1 turns off the local check for experimentation. The server applies the same cap regardless, so it is not a way to submit a denser file.
Why cell counts and controls are strict
Your predictions are compared against held-out data using a differential-expression test. The number of cells per perturbation determines the statistical power of that comparison, and the control cells define its baseline. If either varied between submissions, scores would not be comparable across the leaderboard — so both are fixed: you supply exactly 400 predicted cells per perturbation, and the controls always come from the held-out data.
Common rejections
| Error | Cause |
|---|---|
| "perturbation labels do not match the official list" | Construct ids (ADNP-1) instead of gene symbols, or a perturbation missing from one context. The error names the context and the genes. |
| "N perturbation(s) have the wrong number of cells" | A perturbation has more or fewer than the required 400 cells in that context. |
| "Submission values are FRACTIONAL, but submissions must be RAW INTEGER COUNTS" | Usually log-normalized data — submit adata.raw, or the layer you normalized from. If your model genuinely emits continuous values, round them to whole counts first. |
| "Submission contains NEGATIVE values" | Clip or drop negative predictions. Counts cannot be below zero. |
| "Cell count N exceeds the maximum of 400000" | The whole submission is over the cap — check you haven't concatenated something twice. |
| "Prediction has N stored entries … above the 4,750,000,000 limit scoring can hold" | Too dense, not too big — see the density cap above. Save the matrix sparse, and threshold genes below ~1 CPM if it is still over. Remember explicitly-stored zeros count toward N. |
| "Submission contains N 'non-targeting' control cell(s)" | Drop every control cell before prepping. |
| "Context column 'context' is missing" | Add a context column, or point at yours with --context-col. |
| "Submission is missing predictions for context(s): …" | One file must cover all three contexts. In the final phase the expected labels are D/E/F — this error naming D usually means validation labels were carried over. |
| "Unknown context label(s) in 'context': …" | A label that isn't in the bundle you're predicting for. Same cause as above, in the other direction. |
| "N unexpected perturbation(s) present" | An extra perturbation your submission predicts that pert_counts.csv doesn't list. Reported alongside the mismatch above. |
Two of these come from scoring rather than vcc prep, and you'll only see them if you submitted a file that didn't go through prep: "N perturbation(s) are not in the panel" and "Submission contains unknown context label(s)". They mean the same things as their prep equivalents above.
Full command reference
Every command's flags, defaults and help text: CLI reference. That page is generated from vcc --help, so it always matches the version you have installed — and vcc <command> --help in your own terminal shows the same thing.
Troubleshooting
| Symptom | Fix |
|---|---|
command not found: vcc | Run uv tool update-shell, open a new terminal. Don't reinstall. |
vcc login refuses to store the key | Expected in a remote or headless environment with no OS keychain. Use export VCC_TOKEN=<key> instead (or --store-plaintext if you accept a 0600 plaintext file). |
| Lost your API key | Generate a new one from Account → Credentials. The old key is revoked automatically. |
whoami says "NOT able to submit yet" | The output lists the blocker — most commonly registration/verification is still pending. |
Scorer rejects a vcc sample file with "perturbation mismatch" | The sample was built without -p pert_counts.csv. Rebuild with -p. |
| Download or upload interrupted | Downloads resume automatically on re-run; uploads resume with vcc submit <file> --resume. |
| HTTP 409 "already has a submission in progress" | Your team already has a non-terminal submission. Wait for it (vcc status <id> --wait). This guard is intentional. |
The -p flag means different things in sample and prep
| Command | -p means |
|---|---|
vcc sample -p <file> | --perts — the perturbation list |
vcc prep -p <name> | --pert-col — a column name |
Copying -p pert_counts.csv from the sample step into prep does not error: it quietly sets the perturbation column to the string "pert_counts.csv" and then fails with a confusing "column not found". In prep the flag is --perts, with no short form.
Still stuck? Ask in the Discord community or reach out to us for support.