MARKET MAP / RECHECKED 2026-09-02

The pieces exist.The rights-cleared conversion chain remains unproven.

Session archives, evaluation platforms, RL environments and expert-data suppliers are documented product categories and active programs. Within the official sources reviewed as of 2026-09-02, we did not identify one public service documenting the entire combination of heterogeneous local coding-agent capture, user-scoped rights, environment reconstruction, multi-model reruns, artifact-bound verification, contributor compensation and buyer delivery. That is a bounded search result, not a market-wide proof of absence.

FIELD NOTES / DEEP DIVES

RLDR / 학습 소스

AI가 배우고 일하는 방식을 영상과 글로 살펴봅니다

RLDR 영상의 논지와 예시, 작동 원리와 한계를 실제 장면과 원문 출처를 따라 읽습니다.

2026.09.17 / TASK NOTEBOOK

에이전트 task 하나를 직접 만들어보면

Harbor, Prime Verifiers, Inspect에서 같은 원장 정리 작업을 만든다. 완성된 파일 구조, 만드는 순서, LLM이 들어가는 자리를 함께 연다.

2026.09.17 / CONTINUAL LEARNING

AI를 고친 시간은 누구의 자산이 되는가

Trajectory와 기업별 continual learning. 직원의 수정이 다음 모델의 능력으로 남으려면.

2026.09.17 / RL ENVIRONMENTS

Mechanize, 다음 태스크에 남는 연구

RL 환경을 만드는 작은 팀의 값어치. 납품한 태스크와 다음 제작에 남는 연구를 함께 본다.

01 / APPLIED COMPUTE

The company owns the learning loop, not just the model

Five talks and primary sources on the business, technology and future of Specific Intelligence.

02 / TERMINAL-BENCH

How hard work becomes a measurable agent task

A field study of 328 proposals, 624 task PRs, and the anti-cheat evolution of rs-archive-clone.

BOUNDED HYPOTHESIS

Datafooding's whitespace hypothesis is not another log viewer or another environment marketplace. The target process keeps private sessions intact, selects one authorized task, reconstructs a bounded environment, reruns multiple model/harness conditions, and would issue an artifact-bound verification receipt for a future buyer review.

01 / Archive & browse

CASS

Local indexing and search across the coding-agent providers named in its public README, reviewed 2026-09-02.

VS Code / GitHub

Copilot sessions can sync private-by-default and appear in Chronicle. VS Code separately discovers supported local Claude Code and Codex sessions; the public docs do not say those external sessions are synced to GitHub.

02 / Replay & verify

Kitaru

Imports or records agent traces, re-executes actual agent code with recorded tool responses, and compares model, prompt or code changes.

Harbor

Open-source framework for evaluating agents, creating environments and generating rollouts for RL optimization.

HUD

SDK and platform for defining agent evaluations, environments and verifiers, with parallel task runs.

Inspect AI

Open evaluation framework with datasets, solvers, scorers, sandboxing, parallel execution, and external-agent integrations including Claude Code and Codex CLI.

Trajectory

Company material describes instrumenting real product usage, building governed training data and operating a continual model-improvement loop. It is a precedent for the company product category, not evidence that Datafooding already has feature parity.

Applied Compute

Company material describes custom model training, deployment and iteration around a customer's data and evaluation harnesses. It is a market precedent for private-model services, not a shipped Datafooding capability.

Arize Phoenix

Open-source tracing, human or code-based evaluation, versioned datasets, prompt and model comparisons, span replay, and experiments over the same inputs.

BrowserGym

Open, extensible web-agent environment and benchmark framework spanning WebArena, WorkArena, AssistantBench and other task suites.

OSWorld / AppWorld

Two environment precedents outside coding sessions: OSWorld runs agents in real desktop applications; AppWorld provides controllable API-driven apps and execution-scored interactive coding tasks.

Terminal-Bench v4.0.0

v4.0.0, published 2026-08-26, is the latest GitHub release observed in this review. An open v4.1 issue proposes verified Harbor Hub runs for supported-agent submissions; it is not shipped policy.

03 / Supply programs

DataVendor

Public pages document beta dataset listings, Stripe checkout, sample purchase and managed Harbor or HUD delivery. We found no independent evidence of completed sales or buyer acceptance in the reviewed sources.

Scale Data Partnerships

A direct rights-governed data-supply competitor for enterprise workflows: its public material describes structured and de-identified records, per-package approval, audit trails and on-prem or VPC options. It is not coding-session-specific.

Surge AI

Public product material describes RL environments and agents, rubrics and verifiers, RLHF, SFT and human evaluation.

Turing Frontier AI

Public product material describes expert-authored datasets, RL environments, trajectory capture, and custom or off-the-shelf datasets available through a sample or engagement request.

Slesh

Public company material describes human-verified browser trajectories collected from real tasks for training.

RSI Bench

A public benchmark contribution effort: each task has instructions, a baseline, an environment and a verifier. It lists $2,000 per accepted task for the initial cohort of 50, and says benchmark data must not enter training corpora. This is RSI Bench's program, not a Datafooding payment or buyer offer.

EXACT BUNDLE TEST

The exact chain remains a dated, bounded hypothesis.

01

Archive competitors

Codecast, Entire, CASS, AgentLens and VS Code/GitHub already cover meaningful parts of ingestion, sync, search and browsing. We do not assume archive UX alone is a moat.

02

Replay competitors

Kitaru already covers trace-backed replay and model, prompt or code-change testing. Harbor, HUD, Inspect, Phoenix, Weave, BrowserGym, OSWorld, AppWorld and Prime cover substantial evaluation, environment, rollout or comparison primitives.

03

Data suppliers

DataVendor, Scale, Surge, Turing and Slesh document adjacent marketplace or supply workflows. Public offers do not prove completed transactions or buyer acceptance.

04

Datafooding whitespace hypothesis

Compile a user-selected session into an independently reproducible, rights-valid comparison unit. Measure yield and accepted-unit cost, not stored bytes.

REAL / SYNTHETIC / BOUNDARY / NOT YET

A useful report must also say what has not happened.

OBSERVED

Controlled owner-Mac receipts demonstrate, for the tested artifacts only, client-encrypted GCS upload, remote-only bit-perfect restore, agent-home archives and source preservation. They do not establish universal safety or continuous production reliability.

OBSERVED PROTOTYPE

A prospective Git snapshot can be paired with a fixed Datafooding harness, model targets, isolated workspaces and comparison-pack metadata.

SYNTHETIC

Current runner tests use fake provider and sandbox implementations. Hosted comparison review has not been proven by a production buyer receipt.

BOUNDARY

Comparison submissions and prototype credits are internal review artifacts. They do not trigger user payment, buyer delivery, training rights or resale.

NOT YET

Historical Claude/Codex/Hermes product-stack replay, independently provisioned verifier attestation and fleet, buyer delivery, external training and redeemable payout. A local HMAC-authenticated registry and mutation controls exist only as bounded pilot evidence.

The first decisive proof is deliberately small: one authorized historical task, two or more models, at least three repeats, an independent operator restore, and one prospective buyer acceptance criterion.