Aman’s Jev lab

See where the judgment lands.

One map looks back at this session. The other finds the context worth bringing into your next task.

Loading recorded Jev results…

The skill is installed. The workflow is the next step.

Bessi’s reference points to the TypeSafe skill already used in this work. It teaches the coding agent how to integrate Jev; it does not itself run a supervisor.

Recommended first integration: selected claim + cited evidence → Jev judgment → human review → a replayable eval case. Code owns access checks and execution. Jev helps assess meaning; tests establish whether a fix works.

Proposed integration, not a deployed trace checker. No new private data transfer or background monitoring is enabled.

New use case: live instruction-drift detection

Josh Rosen’s post introduces Foreman: Jev observes worker activity against repository instructions, while code controls interventions.

Our proposed experiment: observe first, compare flagged moments with human labels, and measure false alarms before enabling steering. A useful next visual is a timeline of requirements, observed behavior, drift judgments and verified recovery.

Reference added September 19 · Not implemented here. Existing heatmaps are unchanged; no live workers are being monitored or interrupted.

Small questions. Visible evidence. Code in control.

What Jev actually evaluated

24 independent Choice questions assess six summarized session stages. Another 40 Noul questions ask whether each context fragment would help each task. These are real, recorded API responses—not fresh calls when you click.

The summaries were prepared by the assistant being assessed, using the visible conversation and prior notes. They are not a raw-trace audit or an independent verdict. Context comes from a previously inspected sample of Traces, local preferences and project history, generalized for this public page.

What the colors don’t prove

Trace cells show probability of “met”; unknown and not-applicable judgments stay distinct. Confidence describes distribution concentration, not correctness. Context cells show probability of relevance—not importance or confidence.

The inclusion threshold is an illustrative, uncalibrated policy. It only changes a preview shortlist. No memories are deleted, no tools are executed, and no API keys or raw private traces are published.

TypeSafe building guide · Grok idea reference

The source post’s speed/cost multipliers and “zero hallucination” claim are not treated as verified benchmarks.