The Dashboard
The dashboard is what you arrive at after the QuickStart: the running container’s control surface, in your browser. It is where you run pipeline steps, inspect outputs, attest that you have looked at them, climb the PROOF reproducibility ladder, push results to GitHub or Overleaf, and (optionally) let an AI coding agent work alongside you.
This page is a tour of every panel.
Layout
The dashboard has a fixed layout:
Top toolbar — container name, active project, the three PROOF level badges, the ? Help button, and the Run, Sync, View, and Admin menus. View → Resource Monitor opens a small on-demand panel with live CPU and memory sparklines and the container’s disk usage (with a warning banner when the disk is nearly full); when a reading is unavailable — Docker unreachable, container stopped — the panel says so rather than showing a stale number.
Left panel — a tabbed panel. For projects with a
project.jsonthe tabs are Main, PROOF, Files, and Logs; for sandbox and toolkit projects (noproject.json) they are Files, Repos, and Logs.Top panels — Two “Viewing Windows” to display plots and files.
Bottom panel(s) — Terminal window(s)/tab(s) for work inside the container.
Beside the project name, three copies of the vaibify badge mark PROOF Levels 1–3 (Self-Consistent, Published, Reproducible). Each lights up when the project attains that level, and the whole dashboard theme shifts color with the highest level attained: pale blue before Level 1, purple at Level 1, green at Level 2, and pink at Level 3. The badge, the logo, and every “attained” mark share the tint, so a glance at any corner of the screen tells you where the project stands.
Terminal
Containerized projects get a shell in the dashboard. The terminal strip opens a session inside the container, with tabs and panes.
It costs one guarantee, and you should know which. A shell can start a process that detaches from the process group vaibify records for it, so vaibify cannot prove such a process has stopped. Rather than claim otherwise, vaibify stops claiming: once a terminal has run in a container, releasing it, handing it to another session, or shutting the hub down reports the container’s quiescence as unproven instead of quiet, and asks you to settle it with:
vaibify reconcile
Nothing is lost silently — a container in that state is quarantined until reconcile proves it, rather than being reported clean.
Getting text out of a pane
A full-screen program — an agent, vim, htop — asks the terminal
for the mouse, and while it holds it your drag belongs to the program
rather than to the text. Dragging selects nothing, and the pane jumps
to its newest line. Three ways out, in the order you will want them:
Select text, in the pane’s tab bar, gives the mouse back to the pane. While it is on the button is filled in, dragging selects, dragging past the bottom edge scrolls further text into reach, and the wheel moves through the scrollback. Turn it off to give the mouse back to the program.
Shift + wheel always scrolls the pane, whether or not a program is holding the mouse, and whether or not the mode above is on.
Copy all puts the pane’s whole scrollback on the clipboard, with wrapped lines rejoined as the program wrote them. It needs no gesture at all, so nothing can take it from you.
Selected text reaches the clipboard on its own; Cmd+C, Ctrl+Shift+C, Ctrl+Insert and right-click all copy it as well.
A project that runs on your machine has no in-dashboard terminal.
There is no container to open a shell inside, and your own terminal is
the same shell with the same authority; the strip names the project’s
directory so you can cd straight to it.
You can always reach a container’s shell outside vaibify:
docker exec -it <container-name> bash
That is outside vaibify’s containment, which is exactly the point: the responsibility for what you start there is visibly yours, not silently vaibify’s.
If an agent CLI is enabled for the project, start it from that shell. For example:
claude --dangerously-skip-permissions
codex
gemini
opencode
cline
openhands
pi
Each CLI receives the same Vaibify context, skills, and vaibify-do
dashboard bridge. Claude’s option skips its per-command prompts; use an
equivalent unattended mode for another provider only when you explicitly
intend it. The container is an isolated sandbox, the agent runs as an
unprivileged user with no sudo, and everything it edits is tracked in git
and hash-pinned in the project manifest. Your protection comes from
verifying results, not from approving each
command — see the Using AI section of the Help panel
and the Security model. The agent can in turn ask the
dashboard to run steps, generate tests, push to GitHub, and so on —
see Agent actions below.
Viewing Window
The Viewing Windows above the terminal strip display plots and ASCII text files in the container. Supported formats include PDF, PNG, SVG, and JPG. In Project mode, the log is displayed in a window.
Moving files in and out
Into a project. Drag files, or whole folders, from your computer onto
the Files tab. Drop onto the drop zone to upload into the folder you
are viewing, or onto a folder row to upload into that folder. A dropped
folder is recreated with its subfolders and its empty folders; any
.git or .vaibify folders inside it are left out and counted in the
summary. A file of the same name is only replaced after you confirm it,
and a name that appeared while the upload was running is reported rather
than overwritten.
There is no fixed size limit: the limit is the space the disks have. Vaibify checks there is room, and says which disk is short, before it sends the first byte. That is an upfront check, not a promise: a disk that fills while the file is being placed stops the upload and leaves the previous file untouched. The drop zone is disabled, with the reason, for a folder uploads cannot go into (git’s own folders, Vaibify’s metadata folder, or anywhere outside the workspace).
Out of a project. Right-click a file and choose Download to this computer, or a folder and choose Download as .tar. Files land in your browser’s downloads folder, on the computer you are sitting at, whether or not Vaibify runs there. A symbolic link inside the project downloads as the file it points to; one that points outside the project is refused and the message names it. A folder’s tar keeps its links as links.
From the terminal. vaibify push and vaibify pull copy the same
way, and paths inside the project may be relative. See
Shell helpers for TAB completion.
Opening the container in VS Code
View > Open in VS Code opens the running container in a new VS Code window, so the window you already have, and any unsaved files in it, are left alone. It needs the Dev Containers extension. The button is hidden in a remote session, because the link names a container that exists only on the remote machine’s Docker daemon.
The first time, two confirmations appear that are not Vaibify’s:
Your browser asks whether to open the
vscode://link in Visual Studio Code. Firefox offers a checkbox to remember the answer.VS Code asks “An external application wants to open … Do you want to open this folder?” This is VS Code’s guard against links from other applications; the dialog has a checkbox to stop asking, and the setting behind it is
security.promptForRemoteFileProtocolHandling. Vaibify does not change your VS Code settings.
Repos panel
The Repos panel is the home tab for sandbox and toolkit projects (the
templates without a project.json). In a project the tab is
hidden, but the panel is one click away: every “Open the Repos panel”
link in the Main tab’s Project block and on the PROOF tab lands there.
It lists the git repositories inside the container with their branch,
dirty status, and push controls.
When you first open a container, repositories already present in the
workspace (cloned by the entrypoint from vaibify.yml) are tracked
automatically. If you clone additional repositories inside the
container, vaibify detects them within a few seconds and prompts you to
Track or Ignore them.
Dirty detection reflects whether you have made source-level
changes. Build artifacts that package managers and compilers leave
behind (Python __pycache__/, C *.o, LaTeX *.aux,
*.egg-info/, and so on) are filtered out, so a freshly installed
repository shows as clean unless you have edited its source files.
Push commits and pushes whatever you have staged in the container —
git add, git commit, git push rolled into a single button. A
secondary Push files… option in the gear menu opens a file picker
for selecting specific files to commit.
The Main tab
The Main tab is the project’s control surface. It contains two top-level collapsible blocks, each with a banner you can click to collapse or expand:
Steps — the per-step work of the project. Level 1 (Self-Consistent) is a per-step property, so this is where Level 1 is earned.
Project — requirements that apply to the project as a whole rather than to any single step. Levels 2 (Published) and 3 (Reproducible) are project-wide, so this is where they are earned.
Both banners carry the same right-aligned strip of status cells as the rows beneath them, so a collapsed block still reports its aggregate state.
The panel header above the blocks holds three buttons: a gear for project settings, a refresh arrow to re-poll remote status, and + to create a new step.
The Steps block
A one-time header row labels the step columns: Run, the warning column (⚠), and L1 | L2 | L3. Each step row then shows, left to right:
Run checkbox — include this step in the next run.
Run light — execution only: hollow gray means the step has not run this session, filled gray means queued, blinking orange means running now, blinking red means running past its runtime limit (see below), solid red means the last run failed, and a quiet pale-blue dot means the last run succeeded. The vaibify check is reserved for attained level cells; the success record (outcome, finish time, durations) lives in the expanded step’s Last run line.
Step label and name — labels are per-type sequential:
A03is the third automated step,I01the first interactive step.Warning column (⚠) — every warning the step carries, consolidated into one glyph; hover it for a plain-English list of reasons and remedies. The color encodes severity, not level: red means something is broken right now (a test failed), orange means pending work or staleness (a script or output changed since verification, an earlier step changed, or a level regressed).
L1 | L2 | L3 cells — the step’s own state at each ladder rung (vocabulary below). Every rung has per-step requirements — Level 2 covers the published copies of this step’s outputs, Level 3 its manifest pinning, determinism, and binaries — and a rung with none shows the muted dash. Clicking a cell opens the step’s detail onto that level’s section.
Steps run from the toolbar’s Run menu: Run Selected Steps, Run All Steps, Clean Outputs, Force Run All (Clean), and Stop This Project’s Run, plus the verification sweeps Verify Outputs, Run All Unit Tests, Verify Dependencies, and Check Files Against Manifest. Steps can be reordered by dragging, and an individual step’s right-click menu offers Run Step, Edit Step, Rename… (which previews and then cascades the rename through the step’s directory, verification marker, manifest paths, and declared paths), Set Runtime Limit…, Run From Here, insertion, and deletion.
Two of those deserve a note, because together they are how you check whether a pipeline still produces the results it recorded.
Clean Outputs deletes every automatic step’s declared output files — data files and figures — and resets each step’s verification marks to untested. It runs nothing afterwards. That is the difference from Force Run All (Clean), which does the same delete and then immediately re-runs everything: the rerun overwrites what was cleared, so you never see the project in its emptied state. Seeing it is the point when you are checking somebody else’s work — an all-gray dashboard is the evidence that what appears next was produced on your machine and not shipped with the repository. Interactive steps keep their outputs (a person made those; nothing here can reproduce them), and anything committed to git can be restored from git.
Check Files Against Manifest re-hashes every file pinned in
MANIFEST.sha256 and reports how many still match. It answers one
question — are these the bytes this project recorded? — and it needs
no container, no network, and no particular PROOF level, so it is
available in host mode at all times. Run it on a fresh clone and it
tells you the published archive is internally consistent. Run it
after cleaning and re-running, and it tells you whether your machine
reproduced those bytes exactly. Both are useful; they are not the
same statement.
Step names and directories
A step’s directory is not a free choice: its final component is
derived from the step name by an obvious formula — remove the
spaces, uppercase each word’s first letter, keep the rest as typed.
“Step Name” becomes StepName; “MCMC 512 Chains” becomes
MCMC512Chains; hyphens pass through (“Grid-Search Sweep” →
Grid-SearchSweep). Names may
contain only letters, digits, spaces, and hyphens, and two steps may
not map to the same directory (compared case-insensitively, because
macOS clones sit on case-insensitive filesystems). Parent folders
are free — analyses/MCMC512Chains is fine; only the last component is
governed.
Creating a step derives the directory automatically; renaming one
moves it (with the full cascade); the generic edit path refuses name
changes so the two can never drift. A step from an older project
whose directory does not match shows a red warning in its ⚠
column, and an Align directories button appears on the Steps
banner: one click migrates every nonconforming step — each move is a
git mv, with markers, manifest entries, and declared paths
following — and reports any step it had to skip (a name containing
now-forbidden characters must be renamed first).
Runtime limit (wall-clock budget)
A running step reports a heartbeat that only proves the runner is alive — a step stuck in an infinite loop keeps that heartbeat beating and, on its own, looks identical to a legitimately long forward-model run. A step’s runtime limit closes that gap: it is a ceiling, in seconds, on how long the step may run before the dashboard flags it. Once the active step passes its limit the run light turns blinking red and its tooltip reads “running longer than its wall-clock budget — may be hung”. This is advisory and never stops the run — an over-limit step keeps executing, because exceeding a declared expectation is not proof of a hang; the flag only tells you where to look.
New projects default to a four-hour limit for every step; existing projects keep whatever they had (no limit unless one was set). The first run of a project shows a one-time notice naming the default and where to change it. Adjust the project-wide default under Settings, or right-click a step and choose Set Runtime Limit… — the dialog prefills a suggestion of twice the step’s last successful runtime, converting “how long should this take?” into “should it take twice as long as last time?”. Zero or blank means the step inherits the project default; a project default of zero disables the feature entirely, so long runs you expect are never mislabeled.
Adding a step
Click + in the panel header to open the step editor. Fill in the step name, working directory, and the commands to run. The editor separates data commands (heavy computation) from plot commands (figure generation), so you can re-run just the plotting after tweaking a script without re-running the simulation. Running a step executes its data commands, then its tests, then its plots.
Interactive steps
An interactive step is one a human drives in a shell. Run in Terminal opens a tab, runs the step’s commands there, and watches for the exit code, so the step’s badge settles when you are done with it.
In a project that runs on your machine there is no terminal in the dashboard (see Terminal), so an interactive step reports the refusal rather than starting something that could never report finishing. Make the step automated, or run its commands in your own shell.
The expanded step view
Clicking a step row expands its detail, which is organized by the reproducibility ladder. At the top sits an optional, expandable Description block — a few sentences on what the step does, written by you or an agent (click the text to edit; agents set the same field through the ordinary step-edit action). There is no separate directory display: renaming a step renames its directory, so the step’s name is its directory. Below come three expandable sections mirroring the banner cells — Level 1 — Self-Consistent, Level 2 — Published, Level 3 — Reproducible — each headed by the same level cell the banner shows, a compact “6/7” count, and an ⓘ that opens a modal listing every requirement of that rung with its live mark and a parenthetical spelling out what the mark means. On first open the step’s target rung — the first level not yet attained — is expanded and the others are collapsed, so the detail opens onto the work the ladder asks for next; your own toggles are remembered after that.
Level 1 is the workbench: the step’s input data, scripts, data analysis commands, output data, plot commands, plot files, test standards, and the Verification section — the step’s own artifacts are exactly its self-consistency surface. It ends with the Run Step button and, just below it, the Last run line (outcome, finish time, wall-clock and CPU durations). File rows carry the per-file marks and remote badges described under Status lights and colors, and clicking a file opens it in a Viewing Window.
Levels 2 and 3 are requirement sections, one row per applicable criterion with a met mark, the offending files when unmet, and either an in-place action (Verify now for the GitHub and Zenodo rows — the same verify the Project block offers) or a pointer to the Project-block section where the project-scoped remedy lives (manifest refresh, environment capture, determinism rules). The rows render the same requirement breakdown the banner cell counts — the two can never disagree. A rung with no requirements for this step says so and its header shows the muted dash.
The Input Data block, between Directory and Scripts, declares the raw files the step consumes that no step produces — for example, observational data committed in the repository. Paths are repo-relative; the + button opens a file picker that browses the container’s project repository (or accepts a typed path). Vaibify watches declared inputs on every poll: a modified input invalidates the step and shows “Input data modified since last run” — the Project is no longer self-consistent until the step re-runs. A step with no raw inputs is declared explicitly with the No input data needed checkbox; a step with neither files nor the checkbox is undeclared and cannot reach Level 1. The Project block’s “Input data declared” row names undeclared steps and offers a one-click bulk declaration for retrofitting an existing Project.
A step that pulls data from a remote source records per-file
provenance (listRemoteData: source URL, retrieval time, content
hash, refreshed after every successful pull). Re-running such a step
when the pulled files already exist asks before overwriting the
canonical committed copy — as a modal in the browser, and as an
actionable refusal for the in-container agent, which must relay the
question to you. Fresh pulls are never auto-committed; review and
commit them through the ordinary canonical flow.
The Verification section at the bottom of the expanded view shows one row per verification axis, each with its state and a timestamp:
Row |
What it records |
|---|---|
Unit Tests |
The combined state of the step’s generated tests; “Last run” is when they last finished, regardless of who ran them. |
Dependencies |
Whether the step’s cross-step inputs are consistent; “Last checked” is the last dependency analysis. |
Your name |
Your own sign-off. Click the row to attest that you have inspected the outputs; “Last updated” is your last attestation. |
Above these rows, plain-English drift notices name exactly which files went stale and why — for example “Tests older than data scripts” or “User verification older than plot files” — so you always know what to re-run or re-inspect. The last run’s outcome and durations live in the Last run line below the Run Step button; the modification times of the step’s data and plot files are shown beside their sections.
The expanded quantitative-tests block additionally carries a Falsification row with a Check test teeth button. It mutation-tests the step’s own Python code against its quantitative tests and records the kill-rate: a statement about the tests’ fault-detection sensitivity — “these tests were shown to notice deliberately injected faults” — never about the result’s accuracy. It is deliberately non-gating (equivalent mutants make a hard pass/fail dishonest) and applies only to deterministic pure-Python steps; a step that shells out to a compiled binary reads not applicable, never green. The record is digest-keyed to the script and its standards, so any edit invalidates it. Runs are on-demand only — cost is roughly mutants × step runtime.
The Unit Tests row expands to the three test categories, with buttons to generate and run them — see Verification.
The Project block
The Project block lists project-scope requirements, grouped into seven collapsible sections:
Section |
What it covers |
|---|---|
Repository |
The Level 1 project-scope requirement: the project lives inside a git repository (its repository). |
Software |
Standalone scientific binaries the project runs, each declared with an expected version and a captured version + SHA-256. |
Artifacts |
The reproducibility envelope files: |
Determinism |
Your declared repeatability rules — how exactly a rerun must match your numbers (random seeding, numeric-library variance). |
Published copies |
The GitHub mirror, Zenodo deposit, Overleaf manuscript, and arXiv submission, with per-file sync state. |
AI |
Five rows. Two count toward Level 2: AI models (which models did the work) and Personal AI Configuration (the researcher’s private host-side agent configuration, accounted for). Three are optional and never count toward a level: Project instructions (the standing agent instructions file), the Prompt Record, and Supervised mode. An optional row shows a dash until it is set up and a green check once it is; it is never red, and anything it needs from you is a ⚠ beside its title, echoed on the collapsed AI heading. The AI Declaration (the researcher’s signed statement) is not a row here — it is a step, and its state lives on the step’s own row. |
Attestation |
The rebuild attestation (Level 3). |
The AI section is where the process provenance lives. Every AI model
used on the project is declared with its vendor, exact model ID, and
dates of use; a closed-weights model passes by declaration, while an
open-weights model additionally declares its weights source and
revision hash. Undeclared is the only failing state. The backend
captures a machine-written stamp (.vaibify/ai_provenance.json) —
the declared models, the SHA-256 of both standing prompt files, the
container’s live network-isolation state, and an explicit trust-base
statement — and folds it into the Level 3 attestation record, so the
archived attestation carries the claim “these models, under these
instructions” alongside the rebuild hashes. The stamp is rewritten on
the next status poll whenever it drifts from the declaration; hand
edits do not survive.
Every section banner and every requirement row inside it carries a status light and an L1 | L2 | L3 level strip: the levels the requirement gates show its state, and the others show a dash. A researcher hunting for Level 2 blockers scans one column.
Expanding a requirement row reveals its file rows (with remote badges), a plain-English status line, one “how to” line, and — where an action exists — a button that performs it in place:
Capture version + SHA and Remove package… on each declared binary, plus Add package… at the bottom of the Software section.
Regenerate now on the manifest, dependency lock, and environment snapshot; Check files against manifest and Check dependencies for on-demand verification.
Generate reproduce.sh to write the one-command reproduction script and pin it in the manifest.
Declare rules / Delete rules… for the determinism declaration (stored directly in
project.json; there is no separate rules file).Configure arXiv… to record the arXiv submission that must match the frozen Overleaf figures (optional — an untracked submission reads “not tracked” and never blocks Level 2).
View Declaration and Edit Declaration on the AI models row (Declare Model when none is declared). The editor holds one card per model, each saved on its own; correcting a vendor or model ID edits that declaration in place rather than adding a second one. Delete asks for confirmation first, because deleting a declaration erases provenance and can drop the project below Level 2 (it is user-only; the in-container agent cannot delete or rename one). When the row is red, it says what is missing.
The Personal AI Configuration row accounts for the fourth layer of the instruction stack that governed the AI’s work: your private host-side agent configuration (global instruction file, personal skills, memory, hooks) — the layers above it being the harness system prompt (declared via the model ID), the vaibify-generated container context, and the project context file. You answer with one of three statuses: none (no personal AI configuration exists), exists — content withheld, or included in the project repository. Answering the question is the Level 2 requirement; disclosure is never required, and “withheld” with nothing else is a fully valid answer. With “withheld” you may optionally add hash commitments: pick a file on your computer and a short label, and vaibify stores only the label, the file’s SHA-256, its byte count, and the date — never the path and never any content. A commitment reveals nothing about the file, but if you later choose to release it, the hash proves the released version is the one that governed the work (it cannot be quietly sanitized after the fact). The hashing runs only from your browser session — the in-container agent is refused, so it can never use the endpoint to probe files on your computer.
Start from template, Import from this computer…, and Adopt repo-root context file when no project context file exists yet. The context file (
.vaibify/AGENTS.md) is the standing instructions the in-container agent reads; the entrypoint symlinks it to provider-recognized root names (CLAUDE.md,AGENTS.md, andGEMINI.md) so any agent tooling finds it. Once it exists, click its file row to view and edit it in place (saves go through a dedicated, path-fixed route; the generic editor still refuses.vaibify/writes). Import is researcher-only — the in-container agent can read and update the file via its own actions but can never pull files from your computer. Edits take effect for the agent’s next session, not a running one.The Prompt Record row is the opt-in transcript recorder. Recording settings (Set up recording while it is off) holds the one decision: whether the in-container agent’s sessions are recorded. When on, the agent’s session transcripts are captured into
.vaibify/promptRecord/every ~30 seconds as redacted transcripts: each capture is scanned (exact vaibify session secrets, the detect-secrets pattern catalog, vendor token prefixes, and a conservative high-entropy check) and every hit becomes a visible[REDACTED: category]marker with a per-category count. Only sessions the agent started inside the project’s folder are captured: a container can hold several projects, so a session started at the workspace root or in another project is left out rather than published into this one; the row says how many were left out, so start the agent from inside the project folder when you want its session recorded. After the first capture, each pass scans only the lines a session has added, and it holds the container’s write lock only to list transcripts and to save results, never while scanning, so a commit or push waits seconds at most. The first pass over a long history still takes minutes; Recording settings says a pass is running and since when, and an agent that asks for a capture meanwhile is told the pass is still running rather than that the host is unreachable.View Record opens the record: the captured sessions with their redaction counts and integrity checks (captures are hash-chained, so editing or removing one breaks the chain loudly, and each session file must still match its capture), the coverage intervals (time the hub was not watching reads as an explicit gap, never implied continuity), and each session as a conversation — your prompts and the agent’s replies as text, tool calls, results and reasoning folded to one line, every redaction highlighted, with a filter that shows only the turns that carry one. Review & Approve (shown until you approve) opens the same view with Approve first capture at its foot; the record counts as a recorded development history only after that approval, and the agent can never approve its own transcript. The record is tamper-evident, not proven complete; enabling it requires
pip install vaibify[replay]on the host.The Supervised mode row is the rung above the Prompt Record, available once the record is on and its first capture approved. When on, every repository change must attribute to a recorded action channel — a pipeline dispatch, an editor save, or a context write — within a 60-second window; changes with no recorded cause become permanent, hash-chained flags (
unattributed-modification), and a repo that changed while the hub was not watching flags anunsupervised-gapon reconnect. Flags are listed on the row with a ⚠ and are never cleared by the tool; editing the flag file breaks its chain loudly. Attribution granularity is honest but coarse: the window and the channel, not the file path. Work done in adocker execshell has no recorded channel at all, so it flags as unattributed — which is the truthful answer, not a gap to paper over.Add AI declaration step if the project has none, and Verify Level 3 reproducibility to launch the full rebuild-and-compare.
The Dockerfile row is guidance-only: the Dockerfile is yours to edit,
and pinning its base image to an exact digest (FROM <image>@sha256:…) is something you — or the in-container agent — do
by hand.
Status lights and colors
The same small vocabulary repeats across step rows, both block banners, and every requirement row. The ? Help panel carries the authoritative legend; this is the summary.
Level cells
The L1 | L2 | L3 cells (and the single L1 cell on step rows) use seven states. The circle fills in as reality does — hollow (nothing exists), gray (material exists), colored (assessed), badge (attained):
Cell |
Meaning |
|---|---|
Hollow gray circle |
Not started — no outputs on disk and no activity at this level yet. |
Gray filled circle |
Unassessed — outputs exist on disk, but no tests, checks, or sign-off have been recorded yet. |
Red circle |
No requirements met. |
Orange circle |
Partially met. |
Vaibify badge (the favicon, theme-tinted) |
Attained — every requirement at this level is met. |
Question mark (?) |
Unknown — GitHub/Zenodo have not been checked recently; refresh remote status to find out. |
Dash (—) |
Not applicable — no requirements at this level for this row. |
The gray states are honest by design: “unassessed” asserts only that the step’s declared outputs exist — hours of compute performed outside the dashboard stay visible as progress — but it never claims verification, and a remote that has never been checked is never shown as passing.
Warning glyphs
Warning glyphs (⚠) are colored by severity, never by level:
Red — broken or failing now: a test failed, a declared file is missing, a requirement check failed.
Orange — pending work or staleness: something changed since the last verification, or a check has gone stale and needs refreshing.
Blue is reserved for purely informational marks, such as the not-tracked-by-git badge. The pencil mark (✎) on a file row means the file changed since its last verified run — re-run the step to refresh it.
File-name styles
Inside expanded rows, a file name rendered in red is itself a diagnosis: upright red means the declared file is missing; red with a dotted underline means it changed since its last test run; red italic means it exists but you have never verified it.
Per-file remote badges
Each file row carries one badge per configured remote (GitHub, Overleaf, Zenodo, arXiv), tinted by that remote’s state:
Badge |
Meaning |
|---|---|
Pale blue |
In sync with the remote. |
Amber |
Local file differs from the last push. |
Red |
Uncommitted local changes. |
Blue |
Not tracked by git (informational). |
Solid muted gray |
Git-ignored — a deliberate |
Faded gray |
Not synced to this remote. |
Only figure formats travel to a manuscript, so the Overleaf and arXiv rows list figure files only.
The PROOF tab
The PROOF tab is the requirements ledger for the reproducibility ladder. A header card names the project’s current level (for example, “Level 1: Self-Consistent”) with a clickable progression strip, followed by three expandable sections — Level 1 — Self-Consistent, Level 2 — Published, Level 3 — Reproducible — each summarizing how many of its requirements are met.
Every requirement row shows a status light, the requirement, what it means, and how to meet it, with a deep link to the surface where the work happens (the Main tab’s blocks or the Repos panel). The tab owns the requirement text; the buttons that do the work live in the Main tab’s Project block. The requirements are:
Level 1: Repository; Every step self-consistent; Project context file (optional — a versioned
.vaibify/AGENTS.mdrecording the agent’s standing instructions; shown with a neutral dash and excluded from the met count when absent, since it never blocks a level).Level 2: GitHub mirror; Zenodo deposit; arXiv manuscript (opt-in — checked only when an arXiv submission is recorded, since posting happens outside vaibify on its own timeline); AI model declared (every model used, with vendor, ID, and dates — closed-weights models pass by declaration); Personal AI Configuration answered (none / exists-but-withheld / included — any answer meets it; disclosure is never required). The AI Declaration’s attestation still gates Level 2 but is homed on its own step row (2026-08-27 ruling — a project-level copy double-counted it); when the step is missing, a ghost row at the bottom of the step list carries the add action.
The AI rows’ deep links land on the specific requirement row in the Project block’s AI section — the link switches tabs, expands the group and row, and scrolls it into view.
Level 3: Manifest complete; Dependency lock; Environment snapshot; Dockerfile pinned; Reproduce script; Determinism declared; Software declared; Envelope published (the envelope files match the copies on the GitHub mirror); Envelope archived (the envelope files are in the Zenodo archive — Zenodo versions are immutable, so this row goes red after any envelope change and comes back at the next published deposit version; that is expected, because Level 3 describes a published release, not the working tree); Rebuild attestation.
The Level 3 section ends with the verification machinery: the Verify Level 3 Reproducibility button (enabled only when the readiness checks pass; the rebuild runs in the container and can take hours), the current Level 3 Attestation card (timestamp, manifest digest, image digest, hashes matched, duration — with a staleness notice if the manifest has changed since), and the Reproduction History table of every attempt.
See Reproducibility for what each envelope artifact contains and how third parties verify it without vaibify.
Admin > Environment Info
Every value in this modal is a fact read off the image this session is connected to, or the word unknown. Nothing falls back to the vaibify installed on your machine: the installed recipe describes the image vaibify would build today, which is a different question from what you are running, and answering the second with the first would state an environment your container never had.
Row |
What it is |
|---|---|
Digest / Image ID |
The image’s content identity. This is what Level 3 pins and what |
Architecture |
The platform the image was built for. |
Recipe fingerprint |
A hash over the build inputs vaibify controls — the Dockerfile, package lists, entrypoint, your |
Archive epoch |
The date of the package archive this image’s C compiler and libc came from. |
Inside the container |
Versions probed from within the container itself, not from the host. |
The archive epoch is the row worth understanding. It is pinned, so every image built from this recipe gets the same compiler whatever day it is built, and it moves only when a maintainer deliberately moves it — a change that asks you to re-run and re-verify. See The toolchain epoch for what moving it means.
The epoch covers the compiler only. Editors, LaTeX, the Python
interpreter and pip packages track their upstreams and can differ
between two builds of the same recipe. Pinning those for a particular
result is what the project’s requirements.lock is for — which is the
honest division of labor: the image freezes what compiles your
binaries, and the lock file freezes what runs your analyses.
unknown is a real answer here, not a blank. An image built before vaibify recorded the epoch carries no label, and the modal says so rather than guessing.
The hub’s help and diagnosis
The environment hub (the landing page) has a ? beside the ⧉ New window icon. It opens three short blocks: creating an environment (starting with +), a legend of the tile marks drawn with the real glyphs, and troubleshooting tips such as the one-environment-per-tab rule.
Every failure toast on the hub ends with Click to run a diagnosis.
The click runs the host-scope checks of vaibify doctor through the
hub and shows each finding with its level, message, remedy and
command — the same report the terminal prints, so the two cannot
disagree. Docker’s own error text, when vaibify recognizes it, is
translated into a sentence about your project (an image that has not
been built, a port another program holds, a container name left from
an earlier start, a full disk) with Docker’s words kept in
parentheses as evidence.
The Project Hub, the list of Projects inside an opened environment, opens with a line above Available Projects: that names the environment it lists (Environment: <name>, the name its tile shows, or Environment: this computer for a project that lives on the host rather than in a container). It has the same ? with three folding blocks: using the Project Hub, a legend, and troubleshooting. Its failures carry the same diagnosis click, its automatic list refresh says when it could not refresh, and a New Project whose directory or name the server refuses says why in the wizard itself.
The Help panel
The ? button beside the project name opens the Help panel. It contains:
A link to the full online documentation.
Using AI — how to start the AI coding assistant from a shell inside the container (
claude --dangerously-skip-permissions) and why skipping per-command permission prompts is the intended, safe mode inside the sandbox: the container isolates the agent from your host, every edit is tracked in git and hash-pinned, and a full rebuild ultimately checks the analysis — the PROOF Level 3 posture.The Legend — the symbol key, in four divisions matching the dashboard’s surfaces: Steps (run checkbox, run light, warning column, per-file marks), Project (requirement-row marks and the Level 2/Level 3 warning catalog), Level status lights (the L1 | L2 | L3 cell vocabulary), and Files and remotes (the per-file badges and red file-name styles).
The legend is generated from the same catalog the dashboard renders from, so it cannot drift from the glyphs you actually see. Status itself is deliberately not in the panel — status lives on the banners and the PROOF tab.
Verification
Steps are verified three ways: 1) unit tests, 2) dependency checks (if applicable), and 3) user attestation. These three controls are displayed in each step’s expanded view.
The Unit Tests row is expandable to show detailed information about the step’s unit tests, including generating and running them. Three categories of unit tests exist:
Integrity tests (
test_integrity.py) — output files exist, are non-empty, load in their expected format, have the correct shape, and contain no NaN or infinity values.Qualitative tests (
test_qualitative.py) — column names, JSON keys, parameter names, and other categorical content match expectations.Quantitative tests (
test_quantitative.pyplusquantitative_standards.json) — numerical output values match stored benchmarks at full double precision, with configurable relative and absolute tolerances.
Test generation is deterministic by default: a Python introspection script runs inside the container, reads each data file, and writes the tests mechanically. No language model is involved on the default path. An LLM-based path is available as a fallback for formats the introspection script cannot read.
See Supported Data Formats for the full list of file types the test generator can read.
vaibify monitors the steps for dependency violations, such as a dependent step
not being fully verified or a dependent file being created after a subsequent step was marked verified.
Finally, clicking the row that carries your name records your own assessment — the human attestation that no test can substitute for.
Publishing and remote sync
The toolbar’s Sync menu holds the publication actions:
Push to GitHub — commit and push the repository to its configured remote.
Push to Overleaf — sync figures and any selected files to the configured Overleaf project.
Archive to Zenodo — upload outputs and receive a DOI.
Configure arXiv… — record the arXiv submission that must match the frozen figures.
Verify Reproducibility — open the remote-verification panel described below.
Credentials for these services are resolved from your host’s keychain
at request time. They are never written into the container or into
vaibify.yml. See External services for the
per-service integration architecture.
Per-file sync state is always visible as the remote badges on file rows, and each remote has a requirement row under Published copies in the Project block. Every one of those rows carries a Verify now button that runs the authoritative remote comparison in place, and a successful Overleaf push re-verifies its row automatically — the row reports the last verification, so the action that refreshes it is always one click away.
The Verify Reproducibility panel
Sync → Verify Reproducibility opens a panel with one row per configured remote (GitHub, Overleaf, Zenodo). Each row shows the same four pieces of information:
Field |
Meaning |
|---|---|
Status pill |
Green / yellow / red, semantics below. |
Summary |
|
Last verified |
Age of the most recent authoritative SHA-256 verify (e.g. “12m ago”). Empty when the remote has never been authoritatively verified. |
Re-verify |
A button that runs an authoritative SHA-256 verify against the remote’s current bytes (downloads the files, recomputes hashes, and compares them against the declared project files as they exist on disk right now — never against |
Pill semantics:
Green — the most recent SHA-256 authoritative verify reported every file matching the manifest.
Yellow — never verified, or drift suspected since the last authoritative verify (the remote’s cheap-poll change-detection layer fired). The remote may or may not actually be out of sync; click Re-verify to find out.
Red — an authoritative SHA-256 verify confirmed at least one file’s hash does not match
MANIFEST.sha256.
A scheduled background loop re-verifies every configured remote on a configurable cadence (default 6 hours), so the panel reflects recently-validated state even if the user never clicks Re-verify.
Hash-aware staleness
Status marks distinguish content drift from cosmetic mtime drift.
After a fresh clone, file mtimes are reset to checkout time, which
historically caused every step to render as stale even though the
bytes had not changed. The dashboard consults the per-file SHA-256
recorded in the test marker before declaring a file stale: a
post-clone or post-touch mtime bump with matching hashes is treated
as content-clean, so the warning column and the ✎ file marks reflect
what is actually on disk, not what a tool merely touched.
Agent actions
When an AI coding agent is running inside the container (Claude Code, Codex, or Gemini), it can ask the dashboard to perform named operations on the user’s behalf. These agent actions are the bridge between the agent’s text-only world and the dashboard’s verified state. This scheme enforces deterministic behavior.
Every state-changing operation in the dashboard — running a step, generating tests, pushing to GitHub, archiving to Zenodo — is registered in a single catalog. Each action carries a stable name, the arguments it accepts, and the verification it triggers when it finishes. The agent never invents an action; it picks one from the catalog or it falls back to plain shell commands.
Shipped agent skills
The container also ships ready-made skills — task recipes the agent loads on demand — installed into the agent’s skills directory at container start (edit or delete your container’s copies freely; an image rebuild refreshes them):
session-budget — keeps long autonomous runs alive across Claude session-usage limits: commit-per-work-unit checkpointing with a running resume note as the primary defense, a conservative usage reading (
claude-monitor, documented as an account-wide lower bound) as the secondary one, and a pause-until-reset mechanic for the 5-hour window. Default pause threshold is 95%; override it by saying so in the task prompt.read-arxiv — token-efficient paper reading: fetch the arXiv e-print TeX source instead of the PDF (far fewer tokens, and figure captions arrive as searchable text), read selectively, record the version read, and fall back to the PDF only when no source exists.
proof-ladder — the ordered L1→L2→L3 walkthrough for raising or auditing a project’s reproducibility level, with the known audit traps codified (
iProofLevelis the only authoritative signal; marker hashes are git blob SHA-1s; publication is user-only).create-pipeline-step — the five-phase protocol for authoring a fully wired step, centered on the
{StepNN.varname}cross-step token contract.vaibify-doc-map — a question→(doc, section) table so the agent reads the right 30 lines of vaibify’s own docs (staged in-container at
/usr/share/vaibify/docs) instead of a whole file.diagnose-failed-run — a triage tree over the read-only
get-pipeline-stateandget-host-log-tailactions for a dead or stuck run.read-manuscript — pull the project’s own Overleaf manuscript (via the
pull-manuscriptaction) into a git-ignored scratch copy and read it, rather than answering from memory.reproducible-analysis — answer any quantitative or statistical question by writing a saved script — never a throwaway one-liner, heredoc, or REPL session — structured so it can become a pipeline step and its number can be regenerated.
running-steps — run pipeline steps through
vaibify-do, never by launching a step’s script directly in a shell: a direct launch is invisible to the dashboard, which can only show runs it is told about.
Moving the ladder and step-authoring walkthroughs into on-demand
skills also slims the always-loaded container CLAUDE.md from ~470 to
~170 lines — a direct per-session token saving, with the
safety-critical rules (authoritative level signal, user-only
publication, the token contract) kept inline.
From inside a shell in the container, you can list the available actions:
vaibify-do --list
Or describe a single action’s arguments:
vaibify-do --describe run-step
Or invoke one directly (the agent does this for you):
vaibify-do run-step A03
Step labels (A03, I01) come from the Steps block — labels are
per-type sequential, so A03 is the third automated step and
I01 is the first interactive step. The dashboard updates as the
action runs; if it produces new files, the affected step’s warning
column and level cell react automatically.
Why this matters
Without the catalog, an agent that wants to “run unit tests on step A03” would have to guess at HTTP endpoints or shell out blindly, and the dashboard would silently drift out of sync with what the agent actually did. The catalog makes the agent’s intentions explicit and verifiable: every action it takes is one the user could have taken through the dashboard, and every action triggers the same verification state machine.
Reproduce a published project
The + button’s Add Environment dialog offers a third kind card beside Container and This machine: Reproduce a published project. It is not a place work runs. It re-runs somebody else’s finished result once, in a shadow container, and leaves no tile behind.
The card opens a dialog with four stages, in this order:
Source. A clone URL (
https://,ssh://oruser@host:path) or the path of a clean clone under your home directory. Stage clones one commit and validates it as reproduction-ready (the six staging rules, not the author’s Level 3 gate). A refusal names the rule and the file. A repository holding several workflows offers a selector rather than picking one by sort order.Confirmation. The project and selected workflow, the resolved commit, the pinned image, the required platform, whether an image deposit is on record, whether this daemon’s architecture matches the pinned build (with an emulation checkbox, off by default, when it does not), and the three links the image will be tried through, in order. Nothing has been pulled or run yet, and nothing is until you click Run; Not now sends nothing.
Progress. The dialog polls only while the hub reports the job live: staging, validating, pulling, downloading (with bytes), loading, running (with the step), comparing, tearing down. It stops the moment the job settles, and it also stops if this hub no longer holds the job – jobs live only as long as the hub that started them, so a restart loses one, and the dialog says so instead of waiting for an answer that will never come. Hide closes the dialog without stopping the job; it reopens on settle.
Result. The verdict – reproduced, reproduced under emulation, diverged, or no verdict with the reason – then every pinned file by name: the ones that diverged, the ones carried in unchanged because a human step produced them, and the ones re-derived byte-identically. Beside them: the three platform facts, where the image came from, the deposit re-check, and a link to the report written under
~/.vaibify/reproductions/reports/.
Dismissing the confirmation deletes the staged clone rather than leaving it: a staged snapshot is a whole repository on your disk, and one nobody is going to run is waste. A staged job you neither run nor dismiss is expired by the hub after half an hour, and the hub holds a small fixed number of staged snapshots at once – stage another and it says so rather than filling the disk.
The report is yours, not the author’s attestation: nothing is written
into the project, the dialog offers no publish, deposit or attest
action, and the shadow container is destroyed with proof when the
comparison is made. The command-line equivalent is
vaibify reproduce --from <source> --rerun; both drive the same
staging, acquisition and shadow lane. See
Reproducibility.
Hub mode
The hub is a separate page from the dashboard, but you visit it every time you launch vaibify with no subcommand:
vaibify
The hub lists every container vaibify knows about on the host, with
quick-launch buttons to start, stop, or open the dashboard for any of
them. From here you can also create a new container (the setup wizard
from the QuickStart) or add an existing project that
already has a vaibify.yml.
Two ways to remove an environment
Each tile’s kebab menu (⋮) ends with two removals, and only one of them is reversible.
Remove from list un-registers the project and touches nothing else. The container, the volumes, the image and every file stay exactly where they are; re-adding the project’s directory brings the tile back with its work intact. This is the right choice for tidying the hub.
Delete environment… is permanent. It removes the container, the workspace volume (every file inside the container you have not pulled out to the host), the credentials volume (any tokens in the container’s keyring), every Docker image built or obtained for the project, and the registry entry. Your project directory on the host, its git history, and anything you have already pushed or published are not touched.
Because it cannot be undone, confirming it means typing
permanently delete <name> exactly — the phrase carries the name, so a
menu opened on the wrong tile cannot be confirmed from memory. The
server validates the same phrase, so the dialog is a courtesy rather
than the gate. An environment that is open in another browser session,
locked by another vaibify process, or holding unsettled operations is
refused rather than deleted, and an in-container agent can never invoke
it.
If any part fails — an image the daemon will not remove, say — the environment stays listed and the dashboard reports what did and did not go. A tile removed over a half-finished deletion would leave bytes on the disk with nothing pointing at them.
Host (uncontained) projects offer only Remove from list. They own no container, no volume and no image, so there is nothing else vaibify could delete without deleting your own directory, which no dashboard button does.
One browser session per container
Each container managed by vaibify can be open in only one browser session at a time. When you claim a container, the server mints a private lease for that tab; a different tab — even another tab of the same browser on the same hub — that tries to open the same container is refused with “In use in another browser session” and its tile renders grayed out. This holds across two tabs of one browser, two browsers, and two hubs alike; it is the single owner-of-record model described in the architecture reference. It is an operational guarantee for honest use behind vaibify’s loopback trust boundary, not a defense against a hostile in-page script.
Reloading the owning tab is safe: its lease lives in sessionStorage,
so the refreshed tab re-asserts the same ownership and is never locked
out of its own container.
A browser that cannot see a tab (a hidden tab, a locked screen, a closed laptop lid) may pause it for longer than vaibify waits for a sign of life, and vaibify then gives the container up. When you come back, the page claims it again by itself and carries on; you are told only if another session has taken it in the meantime, in which case the message says so.
An abandoned session does not hold a container forever. A hub or
viewer left with no connected tab and nothing running self-retires
after an idle timeout (see
Configuration),
freeing its container. Ownership is otherwise released when a brief
disconnect’s grace window expires with no reconnect — never while a
pipeline is still running. Closing the tab does not itself release
anything: the browser fires the same signal on a reload, so treating it
as release intent would drop a running container every time you
refreshed the page.
The hub re-polls availability every few seconds, so a freed container
un-grays on its own without a page reload. You can also list and stop
live sessions from the host with vaibify sessions (see the
CLI Reference).
The New vaibify window icon (⧉) in the hub, the project picker, and the dashboard’s Admin menu opens a fresh vaibify session in a new browser tab — useful for working on two different projects side by side. Each window claims its own containers; it is not a way to open the same container twice.
One container per browser session
The rule also runs the other way: one tab holds at most one container. Starting a container from a tile’s ⋮ menu claims it for that tab before you open it, and the tile then reads held by this tab. Starting or opening a second container from the same tab is refused with “This browser session already holds container …”, and the refusal offers to release the first one and continue. You can also release it yourself from the tile’s ⋮ menu (Release), and returning to the container list from an open dashboard releases in the same way, before anything on screen is torn down. Releasing drops only the tab’s hold: the container keeps running, and a release is refused while a pipeline run or an agent is live in it. When it is refused you stay in the dashboard, with the reason on screen, and your tab keeps its hold and its connection; when the hub does not say whether it released, vaibify asks the hub’s own list before it decides. To work in two containers at once, open a second vaibify window.