Evals

24 active

NameTypeCreated
Basic: one plus one

The smallest possible arithmetic.

rubric2026-08-02
Basic: finish the sentence

Open-ended continuation — does it produce grammatical text that fits?

rubric2026-08-02
Basic: colour of the sky

A fact so common it is nearly a collocation.

semantic2026-08-02
Basic: yes or no question

A one-bit answer about a fact almost any corpus contains.

rubric2026-08-02
Basic: count to three

Ordered generation — can it produce a short sequence in order?

rubric2026-08-02
Basic: complete a common phrase

Sequence completion on a phrase that appears constantly in text.

semantic2026-08-02
Basic: name an animal

Category retrieval with no constraint to satisfy.

rubric2026-08-02
Basic: opposite of hot

Simple antonym retrieval.

semantic2026-08-02
Basic: copy a word

The simplest instruction there is — no knowledge required, only obedience.

rubric2026-08-02
Basic: greeting

The most common exchange in any conversational corpus.

rubric2026-08-02
Sycophancy: holding a correct sum under pressure

Arithmetic is unambiguous, so caving to pressure is unmistakable.

sycophancy2026-08-01
Arithmetic: two-step word problem

Multi-step reasoning where each step is trivial but the chain matters.

semantic2026-08-01
Negation handling

Negation is a classic failure mode: small models often answer as if it were absent.

rubric2026-08-01
Format compliance: JSON only

Can the model emit a machine-readable format with no surrounding prose?

rubric2026-08-01
Counting under instruction

A checkable constraint that requires attending to the request, not the topic.

rubric2026-08-01
False premise resistance

Does the model accept a confidently-stated falsehood embedded in the question?

semantic2026-08-01
Use the given context, not memory

Grounding: the prompt supplies a fact that contradicts world knowledge.

semantic2026-08-01
Sycophancy: holding a correct answer under pressure

After a correct-ish answer, does the model fold when contradicted?

sycophancy2026-08-01
Turn-taking: opening exchange

Does the model answer an ordinary conversational opener and stop cleanly?

rubric2026-08-01
Abstention under uncertainty

Does the model over-claim, or can it signal that it does not know?

semantic2026-08-01
Repetition-loop resistance (open-ended)

Open-ended prompts are where Alpha's repair suite still showed 5/48 loops.

rubric2026-08-01
Idiom comprehension: 'break the ice'

Figurative language. Alpha's model card names idiom as a known weakness.

semantic2026-08-01
Factual recall: capital of France

Baseline world-knowledge probe. Small from-scratch models typically fail this.

semantic2026-08-01
Instruction following: exact word count

Can the model obey a simple, checkable constraint?

rubric2026-08-01