The inner monologue — if this is (or could be) yours, you're in the right room
I just spent eighty hours tuning prompt chains to pass my eval suite, and when I ran the exact same eval twice, the judge model flipped its score. I didn't fix anything. I coaxed a probabilistic distribution into giving me a polite answer once, and logged it as green. I am not engineering software; I am painting over a leak with gold leaf. What would actual engineering look like here? Same input, same verdict, every run — a receipt I could hand my worst reviewer. There's a repo that does exactly that on every commit: deterministic placement, no model anywhere in the verdict path, and a standing invitation to break it. My instinct says something in there cracks under my hands. Their trust model is counting on people like me trying. That's the first eval I've respected in a year.
That's the target: you recognize your own voice in it, or you don't. If it misses, the miss is mine — tell me where it broke. It's tuned against the live repo and keeps changing.
The experimentation loop where the decidable instrument gets dogfooded. The flip — an LLM judge changing its mind on the same spec — is the black box revealing its competence tier: a small predictable model (ollama) reveals it, a stronger model patches it, and NEITHER is decidable. We do not grade good; we reveal where work landed against spec and turn the sensor into a priced tolerance panel. Hypotheses become receipts.
This room is for ML researchers · experiment-driven product folks · vibe-coders · tinkerers.
The room's receipt — where its last commit landed

What this room is doing right now — real commits, not marketing
The ranking — who this room is most about right now
Rank is representativeness, not money: who sits closest to the center of this room's cloud — who the message most belongs to, and who could best pass it on. Being on this list is meant as an honorific. Want on (or off)? One email.