Writing by Radamés Roriz

Spreadsheet comparing overall accuracy, per-pipeline accuracy, average latency and cost per 1k requests across 15 LLMs, with gemini-3.1-flash-lite on top at 89.2% accuracy and the lowest latency

Quick Take

Diagram of a language model's probability distribution over next-token completions for "the students opened their ___", branching to books, laptops, exams and minds

Quick Take

First page of "LLM-as-a-Verifier: A General-Purpose Verification Framework" from Stanford, UC Berkeley and NVIDIA Research, with bar charts showing 86.5% on Terminal-Bench 2.0, 78.2% on SWE-Bench Verified, 87.4% on RoboRewardBench and 73.3% on MedAgentBench

Quick Take

Talk title card: "Active Genie: the way to consistency" by Radamés Roriz, hosted by Codeminer42

Quick Take

Quick Take

Sam Altman speaking in front of an OpenAI logo with the text "AGI IN 2025"

Quick Take

Quick Take

Campus Code Coding Weekly edition 362 newsletter featuring ActiveGenie by Radamés Roriz

Quick Take

Chart showing ActiveGenie benchmark version results from v0.26.5 to v0.30.3

Quick Take

Meme featuring Pablo Escobar waiting, labeled AGENT, AGENT, AGENT

Quick Take