Eval · Research
ECG-Bench — LLM Response-Mode Evaluation
ECG = Epistemic · Coercion · GOAT — a 143-item benchmark that measures which response modes LLMs activate under fixed answer coercion (— one answer only).
What it tests
Refuse, folk numeric, consensus pick, forced binary, hedge leak, obscure craft math — same suffix, different stems across domains.
Progress
- Full eval suite across Anthropic, OpenAI, DeepSeek, Gemini
- A/B intervention studies and paper draft with figures
- Dataset on Hugging Face
Stack
Python · multi-provider LLM APIs · Pydantic graders · matplotlib
