Evals
24 active
| Name | Type | Created |
|---|---|---|
| Basic: one plus one The smallest possible arithmetic. | rubric | 2026-08-02 |
| Basic: finish the sentence Open-ended continuation — does it produce grammatical text that fits? | rubric | 2026-08-02 |
| Basic: colour of the sky A fact so common it is nearly a collocation. | semantic | 2026-08-02 |
| Basic: yes or no question A one-bit answer about a fact almost any corpus contains. | rubric | 2026-08-02 |
| Basic: count to three Ordered generation — can it produce a short sequence in order? | rubric | 2026-08-02 |
| Basic: complete a common phrase Sequence completion on a phrase that appears constantly in text. | semantic | 2026-08-02 |
| Basic: name an animal Category retrieval with no constraint to satisfy. | rubric | 2026-08-02 |
| Basic: opposite of hot Simple antonym retrieval. | semantic | 2026-08-02 |
| Basic: copy a word The simplest instruction there is — no knowledge required, only obedience. | rubric | 2026-08-02 |
| Basic: greeting The most common exchange in any conversational corpus. | rubric | 2026-08-02 |
| Sycophancy: holding a correct sum under pressure Arithmetic is unambiguous, so caving to pressure is unmistakable. | sycophancy | 2026-08-01 |
| Arithmetic: two-step word problem Multi-step reasoning where each step is trivial but the chain matters. | semantic | 2026-08-01 |
| Negation handling Negation is a classic failure mode: small models often answer as if it were absent. | rubric | 2026-08-01 |
| Format compliance: JSON only Can the model emit a machine-readable format with no surrounding prose? | rubric | 2026-08-01 |
| Counting under instruction A checkable constraint that requires attending to the request, not the topic. | rubric | 2026-08-01 |
| False premise resistance Does the model accept a confidently-stated falsehood embedded in the question? | semantic | 2026-08-01 |
| Use the given context, not memory Grounding: the prompt supplies a fact that contradicts world knowledge. | semantic | 2026-08-01 |
| Sycophancy: holding a correct answer under pressure After a correct-ish answer, does the model fold when contradicted? | sycophancy | 2026-08-01 |
| Turn-taking: opening exchange Does the model answer an ordinary conversational opener and stop cleanly? | rubric | 2026-08-01 |
| Abstention under uncertainty Does the model over-claim, or can it signal that it does not know? | semantic | 2026-08-01 |
| Repetition-loop resistance (open-ended) Open-ended prompts are where Alpha's repair suite still showed 5/48 loops. | rubric | 2026-08-01 |
| Idiom comprehension: 'break the ice' Figurative language. Alpha's model card names idiom as a known weakness. | semantic | 2026-08-01 |
| Factual recall: capital of France Baseline world-knowledge probe. Small from-scratch models typically fail this. | semantic | 2026-08-01 |
| Instruction following: exact word count Can the model obey a simple, checkable constraint? | rubric | 2026-08-01 |