BLAH Lens Bundle v1

optional

A stack-neutral format for exposing what happens inside a model, so evals.blah.dev can render a site-by-token view of its internal representations.

The boundary

evals.blah.dev owns the interoperability contract and the visualisation. Each model framework owns faithful execution of its own model. Hugging Face owns immutable distribution of model and lens artifacts.

Nothing in the format or protocol is named after any framework. A bundle produced by PyTorch, JAX, MLX, Candle, Rust, Zig or a bespoke TypeScript engine is indistinguishable to the platform.

This is entirely optional. Models without a lens are registered, evaluated and ranked exactly as before. A bundle only unlocks the model-internals view.

Names

Bundle format:    blah-jacobian-lens
Format version:   1
Runtime protocol: blah-lens-http/1
HF library tag:   blah-jlens

How a model gets a lens

Three related resources: the model checkpoint on Hugging Face, a lens artifact (usually a second repository, because lens artifacts are checkpoint-specific, can be large, and may have several fitting variants), and an execution adapter — either loaded by a worker or exposed by you over HTTP.

your-org/model-name
your-org/model-name-jlens     # recommended

# or inside the model repo
your-org/model-name/evals/jacobian-lens/v1/

Hand this to your coding agent

Run this from the root of your model repository. It inspects your architecture, builds a thin adapter around your native implementation, fits the transports, validates, and prepares the Hugging Face upload — without moving your model into someone else's framework.

or fetch it directly:
curl -L https://evals.blah.dev/lens/agent-prompt.txt -o BLAH_LENS_TASK.md

Verify before you publish

The platform exposes the same validator it runs on import, so nothing needs installing and a bundle cannot pass your check and fail ours. A bundle that fails fatally is rejected rather than stored — a broken lens produces confident, wrong readouts, which is worse than no lens. The response is the full report: every check, its status, and why it failed.

npx @blahai/lens validate ./dist/blah-lens
npx @blahai/lens validate-runtime --url http://localhost:8000 \
  --manifest ./dist/blah-lens/lens-manifest.json

# exit codes: 0 pass, 1 fail, 2 partial

# or with nothing installed:
curl -sX POST https://evals.blah.dev/api/v1/lens/validate \
  -H "Content-Type: application/json" \
  -d '{"lens_repo":"your-org/model-name-jlens","lens_revision":"<sha>"}'

Machine-readable schemas: manifest · protocol

Attach it

Once the bundle is on Hugging Face, attach it to your registered model. The revision must be an immutable commit SHA — a branch name would let the artifact change underneath a published result.

curl -X POST https://evals.blah.dev/api/v1/models/MODEL_ID/lens \
  -H "Authorization: Bearer blah_YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"lens_repo":"your-org/model-name-jlens","lens_revision":"<40-char sha>"}'

The response carries the full conformance report. Your model then has a /models/<id>/lens page. If your runtime declares requires_auth, pass runtime_auth_token in the same body — it is held server-side, sent only to your runtime as a bearer token, and never returned by the read API. Re-attaching without it keeps the token already on file.

Ship example runs in the bundle

Optional but recommended: put finished analyses in examples/<name>.ndjson — one complete protocol stream transcript each, exactly the lines your runtime emits for one analyze call — with an optional examples/<name>.request.json sidecar recording the request that produced it. Import freezes each example into a shareable run page, so your lens page shows real readouts the moment the bundle attaches — before anyone sends a live prompt, and even when execution.mode is precomputed_only and there is no runtime at all.

your-model-jlens/
└── examples/
    ├── capital-of-france.ndjson          # meta, prompt, token…, done
    └── capital-of-france.request.json    # the AnalyzeRequest (optional)

Every line is validated against the protocol schema; a malformed example is skipped and reported in the attach response, never stored. Keep each file under 8 MB; up to 24 are ingested.

Every analysis is a frozen run

A completed analysis is stored under a content-addressed id — sha256(lens fingerprint + canonical request) — announced in the X-Lens-Run-Id response header before the stream begins. Each run has a permanent page at /models/<id>/lens/runs/<run-id> that renders from the stored transcript with no runtime involved, and any two runs — different prompts, checkpoints, or entirely different architectures — compare side by side at /lens/compare.

The frozen run doubles as the cache: a repeated deterministic request (temperature 0) is served straight from it with X-Lens-Cache: hit; pass force: true to recompute. Sampled requests are never cache-served — each becomes its own run. Because the id includes your lens fingerprint, republishing the bundle never serves stale readouts.

curl -sX POST https://evals.blah.dev/api/v1/models/MODEL_ID/lens/analyze \
  -H "Content-Type: application/json" \
  -d '{
    "chat": [{"role":"user","content":"status?"}],   # or "prompt": "…"
    "max_new_tokens": 8,
    "temperature": 0,
    "top_k": 8,
    "modes": ["jacobian","logit"],
    "pinned_token_ids": [1234]    # exact rank/logit at every cell
  }'
# → NDJSON stream; X-Lens-Run-Id / X-Lens-Cache response headers

Analyses too long for an interactive stream go through POST …/lens/deep-run (same request body, up to 256 generated tokens, authenticated): it answers immediately with the run id, executes in the background with live progress on the run's page, and freezes the result there — the browser that requested it can leave. The lens page's Deep run button under Options is the same call.

Execution modes

remote_http — your runtime stays under your control and implements the protocol. Supported today.

huggingface — a platform worker loads your model from Hugging Face. The bundle format is defined; the isolated worker is not deployed yet.

oci — your runtime as a reproducible container, pinned by image digest. Defined; not deployed yet.

precomputed_only — no live prompts; the UI marks live analysis unavailable, and the bundle's examples/ runs are what readers see.

What makes a great lens

Two properties separate a lens page worth staring at from one that merely passes conformance — both are spelled out, with implementation detail, in the agent prompt:

Sites that span the stack. The headline visual is where in the computation a prediction settles. Two adjacent late sites can't tell that story — by then the model has already decided and every column reads "never changed". Publish 8–20 sites at a regular stride, always including the earliest and latest; the UI subsamples deep stacks automatically.

A runtime that streams for real. Conformance measures delivery timing: on a slow analysis the first message must land well before the stream ends, or the runtime fails as buffered. Flush per position (yield the event loop if your compute is synchronous), process the prompt progressively, stop computing when the client disconnects, and emit blank-line keep-alives through long chunks — deep runs abort streams that go silent for 3 minutes, and public proxies kill connections with no bytes for ~100 seconds.

Terminology

The format says site, not "layer" or "residual stream". A transformer's sites happen to be post-block residuals, but a recurrent, state-space or hybrid model's are not, and the grid must not imply they share semantics. Sites are ordered, token-aligned representations — whatever those are for your architecture.