Real session
An approved Codex, Claude Code, Hermes, OpenCode, Kimi Code, or Gemini CLI session.
→MISSION / RL DATA FROM REAL USERS
Your work should train your next model.
Upload real agent sessions. Run one approved task again with other models. Turn verified trajectories into the stack for private models and end-to-end automated research.
Start with the part that works today: one prompt installs encrypted, new-only sync from your Mac to your GCS.
THE COMPOUNDING LOOP
FROM ONE SESSION TO MANY ROLLOUTS
The session is evidence, not the finished product. Datafooding preserves the source, reruns a bounded task across models, and keeps only outcomes that can be traced and verified.
An approved Codex, Claude Code, Hermes, OpenCode, Kimi Code, or Gemini CLI session.
→The same prospective task runs against a frozen environment and verifier.
→Instruction, actions, outcome, grade, provenance, rights, and model receipt stay bound.
→Use the pack for private eval, post-training research, or a rights-cleared lab brief.
THE 60-SECOND HANDOFF
No bucket form. No command assembly. Your coding agent installs the pinned release, creates or verifies your private GCS bucket, and proves automatic sync is running.
It shows the plan before it reads an eligible session.
Install Datafooding on this Mac and finish setup so every next session from the AI agents I choose is automatically backed up to my dedicated GCS bucket.
The outcome I want
- Local session files remain the source of truth and are never deleted.
- Eligible bytes are encrypted on this Mac before upload.
- Only ciphertext is stored in a dedicated bucket in my Google Cloud project.
- Existing history is excluded. Only files created or changed after the new-only baseline may be captured.
- A macOS LaunchAgent keeps capture and sync running every 15 minutes.
Exact release contract
- Version: 0.2.1
- Wheel: https://datafooding.ai/releases/datafooding_agent_vault-0.2.1-py3-none-any.whl
- SHA-256: ba21b23f23a25850adc649a3e48d78c0a8a346193546c8cfbaf63c35dae6247b
- PEP 508 target: datafooding-agent-vault @ https://datafooding.ai/releases/datafooding_agent_vault-0.2.1-py3-none-any.whl#sha256=ba21b23f23a25850adc649a3e48d78c0a8a346193546c8cfbaf63c35dae6247b
Run this workflow
1. Preflight this Mac without reading session content.
- Confirm the OS is macOS.
- Detect only the existence and filesystem metadata of supported stores: Codex, archived Codex, Claude Code, Hermes, Kimi Code, OpenCode, and Gemini CLI.
- Do not open or print prompts, messages, tool results, credentials, tokens, request dumps, hidden reasoning, or session payloads.
2. Install only missing prerequisites from their official distribution.
- Homebrew is the package manager. If it is missing, use the official brew.sh installer and no third-party mirror.
- Install uv with Homebrew when missing.
- Install Google Cloud CLI with the current Homebrew cask: `brew install --cask gcloud-cli`.
- Resolve gcloud with `command -v gcloud`; if the cask is installed but PATH has not refreshed, use `$(brew --prefix)/share/google-cloud-sdk/bin/gcloud` after checking that it is executable.
- Never use a service account, create broad IAM grants, write credentials to the repository, or put secrets in command arguments or chat.
3. Establish the Google Cloud identity interactively.
- Check the active account and project without printing tokens. If login is required, use `gcloud auth login`.
- Never guess a project. If none is active, ask me for the exact project ID.
- Resolve its numeric project number with `gcloud projects describe PROJECT_ID --format='value(projectNumber)'`.
- Propose the dedicated bucket `datafooding-PROJECT_NUMBER`.
4. Ask for one compact approval before creating cloud resources or enabling capture.
Show one block containing:
- the detected agent names, with nothing selected by default;
- the exact account and project ID, redacted where appropriate;
- the proposed bucket;
- for a new bucket, the GCS location I must choose because it is immutable;
- for an existing bucket, its current location and any protection change needed;
- every local prerequisite you installed.
Ask me to reply with the exact sources and `APPROVE`. Treat that answer as consent only for those sources, that project, that bucket, and that location.
5. Create or verify the dedicated bucket, fail closed, and never delete it.
- First run describe. Only an explicit not-found result permits creation. A permission error, timeout, disabled API, billing problem, or ambiguous result must stop with a concrete remediation.
- Create a missing bucket with this exact protection shape:
`gcloud storage buckets create gs://BUCKET --project=PROJECT_ID --location=LOCATION --uniform-bucket-level-access --public-access-prevention --soft-delete-duration=7d --quiet`
- Never set or lock an irreversible retention policy.
- Reuse an existing bucket only when describe proves that its project number matches the selected project and its location matches the approved plan.
- Before continuing, re-describe and verify: exact name, project number, location, uniform bucket-level access enabled, public access prevention enforced, and soft-delete retention greater than zero. If an existing bucket needs a protection update, the approval block must name it before running any update.
6. Install the exact hash-pinned Datafooding release.
Run:
`uv tool install --force 'datafooding-agent-vault @ https://datafooding.ai/releases/datafooding_agent_vault-0.2.1-py3-none-any.whl#sha256=ba21b23f23a25850adc649a3e48d78c0a8a346193546c8cfbaf63c35dae6247b'`
Then require `datafooding --version` to report 0.2.1.
7. Bind only the approved sources from now.
- Run `datafooding quickstart --plan --bucket gs://BUCKET --source SOURCE ...`.
- Verify the content-free plan exactly matches the approved bucket and source list, says `capture_policy: new-only`, `existing_sessions_included: false`, and enables automatic capture.
- Never use `--include-existing`, `archive-home`, or a history migration in this workflow.
- If the plan matches, run the same quickstart command with `--yes`.
8. Trigger and verify local-to-remote sync.
- Run `datafooding sync --capture-enabled`, then `datafooding doctor` and `datafooding status`.
- If the active setup session changes after the baseline, it may become eligible. Let the daemon retry a file that is still changing; do not weaken the stability checks.
- Open `datafooding admin --language en` in a separate terminal because the local admin intentionally stays in the foreground.
9. Call setup complete only when all of these are true.
- Doctor passes.
- Every approved source is enabled with the new-only policy.
- The LaunchAgent is installed, loaded, and bound to the current CLI.
- GCS access and recoverability checks pass.
- Queued ciphertext is zero.
If no post-baseline file has changed yet, report `READY — waiting for the first new session`; do not claim that a session was uploaded.
Failure and rollback rules
- Stop on any identity, project, billing, ownership, location, policy, hash, or health ambiguity. Never retry with wider permissions.
- If quickstart fails after enabling a source that was off before this workflow, disable only that newly enabled source and stop the LaunchAgent started by this workflow. Preserve any pre-existing enabled source and daemon.
- Keep the local vault, Keychain entry, and protected bucket as resumable state. Do not delete source files, vault data, recovery material, or cloud objects.
- Return a short content-free receipt: version, selected source names, bucket protection status, capture policy, LaunchAgent state, queue count, and the exact next remediation if anything is incomplete.Pasting authorizes the pinned local tools. Session capture starts only after you approve the exact new-only plan. Existing history stays out.
THREE PRODUCTS / ONE DATA FLYWHEEL
The same real-user evidence serves a different job at each layer. The status on every product says what can be used now and what still needs a production gate.
Back up Codex, Claude Code, Hermes, OpenCode, Kimi Code, Gemini CLI, and reviewed custom stores without handing us your plaintext archive.
Extract rights-cleared company evidence, convert repeated work into tasks and graders, compare models, and build the smallest learning system that improves held-out work.
A future marketplace for labs to commission benchmark-targeted tasks and compare their own model with other models on matched, rights-cleared real-user work.
PROPOSED CONTRIBUTOR PILOT / NOT LIVE
The first offer is intentionally easy to understand: contribute a reviewed session pack, or create an additional cross-model rollout pack, and receive AI-tool plan value when the pack is accepted.
This is a proposed pilot term, not a live cash offer or guaranteed purchase. 100 GB is a batch unit, not a quality score. Acceptance still requires rights, privacy, provenance, task usefulness, integrity, and verifier review. Today's hosted ‘accepted’ state remains internal and non-redeemable.
THE STACK FOR A PRIVATE MODEL
A session is the source. It becomes learning data only after environment, outcome, verifier, provenance, and rights survive the gates.
Capture approved agent sessions locally, encrypt them before upload, and keep the archive in your GCS.
Run one prospective task through a fixed harness across multiple model targets. This is not exact historical product-stack replay.
Bind the instruction, environment, observable trajectory, outcome, grader, provenance, and rights into a reviewable unit.
Evaluate retrieval, harnesses, adapters, and post-training on held-out work. Keep only the intervention-reducing lift.
Generate the next task from failures, run the model matrix, verify results, and feed evidence back into the learning loop.
START WITH THE SOURCE
Give one setup prompt to the coding agent already running on the Mac that holds your sessions.
Choose the exact agent stores and GCS boundary. No source and no existing history is selected by default.
The agent must prove encryption, storage protection, new-only policy, daemon identity, and an empty queue before saying done.
Keep the raw archive private. Promote only an explicitly approved session into replay, task, review, and future learning.
THE NON-NEGOTIABLE BOUNDARY
Your archive is not automatically training data. Datafooding encrypts eligible bytes before they leave your Mac and stores ciphertext in a bucket you control. Storage, replay, annotation, training, derivative use, and sale require separate rights. The valuable asset is not raw volume; it is a consented trajectory whose environment, intervention, verifier, and outcome can be trusted.