Codecast
Local daemon, multi-agent sync, search and session fork.
MARKET MAP / RECHECKED 2026-09-02
Session archives, evaluation platforms, RL environments and expert-data suppliers are documented product categories and active programs. Within the official sources reviewed as of 2026-09-02, we did not identify one public service documenting the entire combination of heterogeneous local coding-agent capture, user-scoped rights, environment reconstruction, multi-model reruns, artifact-bound verification, contributor compensation and buyer delivery. That is a bounded search result, not a market-wide proof of absence.
FIELD NOTES / DEEP DIVES
RLDR 영상의 논지와 예시, 작동 원리와 한계를 실제 장면과 원문 출처를 따라 읽습니다.
2026.09.17 / TASK NOTEBOOKHarbor, Prime Verifiers, Inspect에서 같은 원장 정리 작업을 만든다. 완성된 파일 구조, 만드는 순서, LLM이 들어가는 자리를 함께 연다.
2026.09.17 / CONTINUAL LEARNINGTrajectory와 기업별 continual learning. 직원의 수정이 다음 모델의 능력으로 남으려면.
2026.09.17 / RL ENVIRONMENTSRL 환경을 만드는 작은 팀의 값어치. 납품한 태스크와 다음 제작에 남는 연구를 함께 본다.
01 / APPLIED COMPUTEFive talks and primary sources on the business, technology and future of Specific Intelligence.
02 / TERMINAL-BENCHA field study of 328 proposals, 624 task PRs, and the anti-cheat evolution of rs-archive-clone.
BOUNDED HYPOTHESIS
Datafooding's whitespace hypothesis is not another log viewer or another environment marketplace. The target process keeps private sessions intact, selects one authorized task, reconstructs a bounded environment, reruns multiple model/harness conditions, and would issue an artifact-bound verification receipt for a future buyer review.
01 / Archive & browse
Local daemon, multi-agent sync, search and session fork.
Git-linked checkpoints for transcripts, tool calls and file changes.
Local indexing and search across the coding-agent providers named in its public README, reviewed 2026-09-02.
Local index and web UI for inspecting and comparing agent traces.
Copilot sessions can sync private-by-default and appear in Chronicle. VS Code separately discovers supported local Claude Code and Codex sessions; the public docs do not say those external sessions are synced to GitHub.
02 / Replay & verify
Imports or records agent traces, re-executes actual agent code with recorded tool responses, and compares model, prompt or code changes.
Open-source framework for evaluating agents, creating environments and generating rollouts for RL optimization.
SDK and platform for defining agent evaluations, environments and verifiers, with parallel task runs.
Open evaluation framework with datasets, solvers, scorers, sandboxing, parallel execution, and external-agent integrations including Claude Code and Codex CLI.
Trace, dataset and experiment workflows with side-by-side comparison, human review, and a local-agent path cloned from a traced run.
Turns production traces and reviewed examples into versioned datasets, evaluations and regression experiments.
Open-source tracing, datasets, experiments and human, model or programmatic scores with API and export surfaces.
Company material describes instrumenting real product usage, building governed training data and operating a continual model-improvement loop. It is a precedent for the company product category, not evidence that Datafooding already has feature parity.
Company material describes custom model training, deployment and iteration around a customer's data and evaluation harnesses. It is a market precedent for private-model services, not a shipped Datafooding capability.
Open-source tracing, human or code-based evaluation, versioned datasets, prompt and model comparisons, span replay, and experiments over the same inputs.
Agent-session tracing, curated evaluation datasets, prompt and model comparison, versioned artifacts, and human annotations.
Community environment hub for RL and evaluation, with hosted runs on Prime-managed infrastructure.
Open, extensible web-agent environment and benchmark framework spanning WebArena, WorkArena, AssistantBench and other task suites.
Two environment precedents outside coding sessions: OSWorld runs agents in real desktop applications; AppWorld provides controllable API-driven apps and execution-scored interactive coding tasks.
v4.0.0, published 2026-08-26, is the latest GitHub release observed in this review. An open v4.1 issue proposes verified Harbor Hub runs for supported-agent submissions; it is not shipped policy.
03 / Supply programs
Public pages document beta dataset listings, Stripe checkout, sample purchase and managed Harbor or HUD delivery. We found no independent evidence of completed sales or buyer acceptance in the reviewed sources.
A direct rights-governed data-supply competitor for enterprise workflows: its public material describes structured and de-identified records, per-package approval, audit trails and on-prem or VPC options. It is not coding-session-specific.
Public product material describes RL environments and agents, rubrics and verifiers, RLHF, SFT and human evaluation.
Public product material describes expert-authored datasets, RL environments, trajectory capture, and custom or off-the-shelf datasets available through a sample or engagement request.
Public company material describes human-verified browser trajectories collected from real tasks for training.
A public benchmark contribution effort: each task has instructions, a baseline, an environment and a verifier. It lists $2,000 per accepted task for the initial cohort of 50, and says benchmark data must not enter training corpora. This is RSI Bench's program, not a Datafooding payment or buyer offer.
EXACT BUNDLE TEST
Codecast, Entire, CASS, AgentLens and VS Code/GitHub already cover meaningful parts of ingestion, sync, search and browsing. We do not assume archive UX alone is a moat.
Kitaru already covers trace-backed replay and model, prompt or code-change testing. Harbor, HUD, Inspect, Phoenix, Weave, BrowserGym, OSWorld, AppWorld and Prime cover substantial evaluation, environment, rollout or comparison primitives.
DataVendor, Scale, Surge, Turing and Slesh document adjacent marketplace or supply workflows. Public offers do not prove completed transactions or buyer acceptance.
Compile a user-selected session into an independently reproducible, rights-valid comparison unit. Measure yield and accepted-unit cost, not stored bytes.
REAL / SYNTHETIC / BOUNDARY / NOT YET
Controlled owner-Mac receipts demonstrate, for the tested artifacts only, client-encrypted GCS upload, remote-only bit-perfect restore, agent-home archives and source preservation. They do not establish universal safety or continuous production reliability.
A prospective Git snapshot can be paired with a fixed Datafooding harness, model targets, isolated workspaces and comparison-pack metadata.
Current runner tests use fake provider and sandbox implementations. Hosted comparison review has not been proven by a production buyer receipt.
Comparison submissions and prototype credits are internal review artifacts. They do not trigger user payment, buyer delivery, training rights or resale.
Historical Claude/Codex/Hermes product-stack replay, independently provisioned verifier attestation and fleet, buyer delivery, external training and redeemable payout. A local HMAC-authenticated registry and mutation controls exist only as bounded pilot evidence.
The first decisive proof is deliberately small: one authorized historical task, two or more models, at least three repeats, an independent operator restore, and one prospective buyer acceptance criterion.