MISSION / RL DATA FROM REAL USERS

Datafooding

Your work should train your next model.

Upload real agent sessions. Run one approved task again with other models. Turn verified trajectories into the stack for private models and end-to-end automated research.

Start with the part that works today: one prompt installs encrypted, new-only sync from your Mac to your GCS.

THE COMPOUNDING LOOP

  • Real user work
  • Matched model rollouts
  • Verifier-bound tasks
  • Private learning loop

FROM ONE SESSION TO MANY ROLLOUTS

Watch real work become learning data.

The session is evidence, not the finished product. Datafooding preserves the source, reruns a bounded task across models, and keeps only outcomes that can be traced and verified.

MATCHED TASK / FIXED-HARNESS PILOTPRIVATE BY DEFAULT
SOURCE

Real session

An approved Codex, Claude Code, Hermes, OpenCode, Kimi Code, or Gemini CLI session.

RERUN

Claude / GPT / Solar

The same prospective task runs against a frozen environment and verifier.

EVIDENCE

Verified trajectory

Instruction, actions, outcome, grade, provenance, rights, and model receipt stay bound.

COMPOUND

Private learning loop

Use the pack for private eval, post-training research, or a rights-cleared lab brief.

ONE SOURCE. THREE PRODUCTS.task → trajectory → outcome → grade

THE 60-SECOND HANDOFF

Give the setup to the agent already on your Mac.

No bucket form. No command assembly. Your coding agent installs the pinned release, creates or verifies your private GCS bucket, and proves automatic sync is running.

  1. 01Copy this setup
  2. 02Paste into Codex or Claude Code
  3. 03Approve the exact sources and bucket once

It shows the plan before it reads an eligible session.

  • Pinned release
  • Private user-owned GCS
  • New-only baseline
  • Verified background sync
Read the exact prompt
Install Datafooding on this Mac and finish setup so every next session from the AI agents I choose is automatically backed up to my dedicated GCS bucket.

The outcome I want
- Local session files remain the source of truth and are never deleted.
- Eligible bytes are encrypted on this Mac before upload.
- Only ciphertext is stored in a dedicated bucket in my Google Cloud project.
- Existing history is excluded. Only files created or changed after the new-only baseline may be captured.
- A macOS LaunchAgent keeps capture and sync running every 15 minutes.

Exact release contract
- Version: 0.2.1
- Wheel: https://datafooding.ai/releases/datafooding_agent_vault-0.2.1-py3-none-any.whl
- SHA-256: ba21b23f23a25850adc649a3e48d78c0a8a346193546c8cfbaf63c35dae6247b
- PEP 508 target: datafooding-agent-vault @ https://datafooding.ai/releases/datafooding_agent_vault-0.2.1-py3-none-any.whl#sha256=ba21b23f23a25850adc649a3e48d78c0a8a346193546c8cfbaf63c35dae6247b

Run this workflow

1. Preflight this Mac without reading session content.
   - Confirm the OS is macOS.
   - Detect only the existence and filesystem metadata of supported stores: Codex, archived Codex, Claude Code, Hermes, Kimi Code, OpenCode, and Gemini CLI.
   - Do not open or print prompts, messages, tool results, credentials, tokens, request dumps, hidden reasoning, or session payloads.

2. Install only missing prerequisites from their official distribution.
   - Homebrew is the package manager. If it is missing, use the official brew.sh installer and no third-party mirror.
   - Install uv with Homebrew when missing.
   - Install Google Cloud CLI with the current Homebrew cask: `brew install --cask gcloud-cli`.
   - Resolve gcloud with `command -v gcloud`; if the cask is installed but PATH has not refreshed, use `$(brew --prefix)/share/google-cloud-sdk/bin/gcloud` after checking that it is executable.
   - Never use a service account, create broad IAM grants, write credentials to the repository, or put secrets in command arguments or chat.

3. Establish the Google Cloud identity interactively.
   - Check the active account and project without printing tokens. If login is required, use `gcloud auth login`.
   - Never guess a project. If none is active, ask me for the exact project ID.
   - Resolve its numeric project number with `gcloud projects describe PROJECT_ID --format='value(projectNumber)'`.
   - Propose the dedicated bucket `datafooding-PROJECT_NUMBER`.

4. Ask for one compact approval before creating cloud resources or enabling capture.
   Show one block containing:
   - the detected agent names, with nothing selected by default;
   - the exact account and project ID, redacted where appropriate;
   - the proposed bucket;
   - for a new bucket, the GCS location I must choose because it is immutable;
   - for an existing bucket, its current location and any protection change needed;
   - every local prerequisite you installed.
   Ask me to reply with the exact sources and `APPROVE`. Treat that answer as consent only for those sources, that project, that bucket, and that location.

5. Create or verify the dedicated bucket, fail closed, and never delete it.
   - First run describe. Only an explicit not-found result permits creation. A permission error, timeout, disabled API, billing problem, or ambiguous result must stop with a concrete remediation.
   - Create a missing bucket with this exact protection shape:
     `gcloud storage buckets create gs://BUCKET --project=PROJECT_ID --location=LOCATION --uniform-bucket-level-access --public-access-prevention --soft-delete-duration=7d --quiet`
   - Never set or lock an irreversible retention policy.
   - Reuse an existing bucket only when describe proves that its project number matches the selected project and its location matches the approved plan.
   - Before continuing, re-describe and verify: exact name, project number, location, uniform bucket-level access enabled, public access prevention enforced, and soft-delete retention greater than zero. If an existing bucket needs a protection update, the approval block must name it before running any update.

6. Install the exact hash-pinned Datafooding release.
   Run:
   `uv tool install --force 'datafooding-agent-vault @ https://datafooding.ai/releases/datafooding_agent_vault-0.2.1-py3-none-any.whl#sha256=ba21b23f23a25850adc649a3e48d78c0a8a346193546c8cfbaf63c35dae6247b'`
   Then require `datafooding --version` to report 0.2.1.

7. Bind only the approved sources from now.
   - Run `datafooding quickstart --plan --bucket gs://BUCKET --source SOURCE ...`.
   - Verify the content-free plan exactly matches the approved bucket and source list, says `capture_policy: new-only`, `existing_sessions_included: false`, and enables automatic capture.
   - Never use `--include-existing`, `archive-home`, or a history migration in this workflow.
   - If the plan matches, run the same quickstart command with `--yes`.

8. Trigger and verify local-to-remote sync.
   - Run `datafooding sync --capture-enabled`, then `datafooding doctor` and `datafooding status`.
   - If the active setup session changes after the baseline, it may become eligible. Let the daemon retry a file that is still changing; do not weaken the stability checks.
   - Open `datafooding admin --language en` in a separate terminal because the local admin intentionally stays in the foreground.

9. Call setup complete only when all of these are true.
   - Doctor passes.
   - Every approved source is enabled with the new-only policy.
   - The LaunchAgent is installed, loaded, and bound to the current CLI.
   - GCS access and recoverability checks pass.
   - Queued ciphertext is zero.
   If no post-baseline file has changed yet, report `READY — waiting for the first new session`; do not claim that a session was uploaded.

Failure and rollback rules
- Stop on any identity, project, billing, ownership, location, policy, hash, or health ambiguity. Never retry with wider permissions.
- If quickstart fails after enabling a source that was off before this workflow, disable only that newly enabled source and stop the LaunchAgent started by this workflow. Preserve any pre-existing enabled source and daemon.
- Keep the local vault, Keychain entry, and protected bucket as resumable state. Do not delete source files, vault data, recovery material, or cloud objects.
- Return a short content-free receipt: version, selected source names, bucket protection status, capture policy, LaunchAgent state, queue count, and the exact next remediation if anything is incomplete.

Pasting authorizes the pinned local tools. Session capture starts only after you approve the exact new-only plan. Existing history stays out.

THREE PRODUCTS / ONE DATA FLYWHEEL

Personal. Company. Lab.

The same real-user evidence serves a different job at each layer. The status on every product says what can be used now and what still needs a production gate.

01 / PERSONALPRIVATE ARCHIVE AVAILABLE

Own the session. Choose the upside.

Back up Codex, Claude Code, Hermes, OpenCode, Kimi Code, Gemini CLI, and reviewed custom stores without handing us your plaintext archive.

  • Encrypted local-to-GCS sync
  • Choose which history can enter review
  • Proposed Session Pack: $200 AI-plan value per accepted 100 GB
  • Proposed Rollout Pack: rerun across models, then submit the accepted pack
02 / COMPANYDESIGN-PARTNER SCOPE

Turn internal work into your model stack.

Extract rights-cleared company evidence, convert repeated work into tasks and graders, compare models, and build the smallest learning system that improves held-out work.

  • Private data and task inventory
  • Environment, grader, and rollout factory
  • Evaluation before RAG, adapter, LoRA, or RLVR
  • End-to-end automated research loop
03 / LABBUYER PLATFORM NOT LIVE

Buy evidence for the capability you need.

A future marketplace for labs to commission benchmark-targeted tasks and compare their own model with other models on matched, rights-cleared real-user work.

  • Target capability and benchmark brief
  • Matched multi-model rollout packs
  • Verifier, provenance, consent, and deletion lineage
  • Curated delivery only after independent release review

PROPOSED CONTRIBUTOR PILOT / NOT LIVE

$200 in AI-plan value per accepted 100 GB pack.

The first offer is intentionally easy to understand: contribute a reviewed session pack, or create an additional cross-model rollout pack, and receive AI-tool plan value when the pack is accepted.

  • Session Pack: accepted rights-cleared session logs
  • Rollout Pack: accepted matched reruns of an approved task
  • Fulfilment target: Claude or equivalent AI-plan value, not a platform token

This is a proposed pilot term, not a live cash offer or guaranteed purchase. 100 GB is a batch unit, not a quality score. Acceptance still requires rights, privacy, provenance, task usefulness, integrity, and verifier review. Today's hosted ‘accepted’ state remains internal and non-redeemable.

THE STACK FOR A PRIVATE MODEL

Real work enters once. Evidence compounds.

A session is the source. It becomes learning data only after environment, outcome, verifier, provenance, and rights survive the gates.

01WORKS TODAY

Session

Capture approved agent sessions locally, encrypt them before upload, and keep the archive in your GCS.

02BOUNDED PROTOTYPE

Rollout

Run one prospective task through a fixed harness across multiple model targets. This is not exact historical product-stack replay.

03LOCAL PROTOTYPE

Task

Bind the instruction, environment, observable trajectory, outcome, grader, provenance, and rights into a reviewable unit.

04VISION / GATED

Private model

Evaluate retrieval, harnesses, adapters, and post-training on held-out work. Keep only the intervention-reducing lift.

05VISION / GATED

Auto research

Generate the next task from failures, run the model matrix, verify results, and feed evidence back into the learning loop.

START WITH THE SOURCE

First, upload your next session.

01

Paste

Give one setup prompt to the coding agent already running on the Mac that holds your sessions.

02

Approve

Choose the exact agent stores and GCS boundary. No source and no existing history is selected by default.

03

Verify

The agent must prove encryption, storage protection, new-only policy, daemon identity, and an empty queue before saying done.

04

Compound

Keep the raw archive private. Promote only an explicitly approved session into replay, task, review, and future learning.

THE NON-NEGOTIABLE BOUNDARY

Private by default. Learning by explicit choice.

Your archive is not automatically training data. Datafooding encrypts eligible bytes before they leave your Mac and stores ciphertext in a bucket you control. Storage, replay, annotation, training, derivative use, and sale require separate rights. The valuable asset is not raw volume; it is a consented trajectory whose environment, intervention, verifier, and outcome can be trusted.