[{"data":1,"prerenderedAt":103},["ShallowReactive",2],{"mdc--wfgsgd-key":3},{"data":4,"body":5},{},{"type":6,"children":7},"root",[8,16,21,34,64,69,74,79,84,89],{"type":9,"tag":10,"props":11,"children":12},"element","p",{},[13],{"type":14,"value":15},"text","The framework around the model is worth more than the frontier model itself.",{"type":9,"tag":10,"props":17,"children":18},{},[19],{"type":14,"value":20},"Researchers from Stanford, UC Berkeley, and NVIDIA Research showed that you can beat GPT-5.5, Claude Opus 4.8, and Gemini 3.1 Pro without upgrading to a better model. Instead, they used a tournament technique where candidate responses face off in 1v1 matchups.",{"type":9,"tag":10,"props":22,"children":23},{},[24,26,32],{"type":14,"value":25},"It's called ",{"type":9,"tag":27,"props":28,"children":29},"strong",{},[30],{"type":14,"value":31},"Probabilistic Pivot Tournament",{"type":14,"value":33},", and it works like this:",{"type":9,"tag":35,"props":36,"children":37},"ol",{},[38,44,49,54,59],{"type":9,"tag":39,"props":40,"children":41},"li",{},[42],{"type":14,"value":43},"The model generates N candidate responses for the same prompt.",{"type":9,"tag":39,"props":45,"children":46},{},[47],{"type":14,"value":48},"A fast initial pass assigns a baseline score to each response.",{"type":9,"tag":39,"props":50,"children":51},{},[52],{"type":14,"value":53},"Based on that ranking, responses split into two groups: the top performers (the pivot group) and the rest.",{"type":9,"tag":39,"props":55,"children":56},{},[57],{"type":14,"value":58},"Each non-pivot is paired against every pivot, while pivots compete among themselves in 1v1 duels.",{"type":9,"tag":39,"props":60,"children":61},{},[62],{"type":14,"value":63},"The duel outcomes create the final ranking, selecting the single best response.",{"type":9,"tag":10,"props":65,"children":66},{},[67],{"type":14,"value":68},"It isn't fast, and it isn't cheap. But it works: 86.5% on Terminal-Bench V2 compared to 84.7% for GPT-5.5, and 78.2% on SWE-Bench Verified compared to 76.8% for Opus 4.5.",{"type":9,"tag":10,"props":70,"children":71},{},[72],{"type":14,"value":73},"This is what changes everything: top performance on a task no longer relies solely on the frontier model, but on the harness and engineering around it.",{"type":9,"tag":10,"props":75,"children":76},{},[77],{"type":14,"value":78},"If you have followed my work for a while, you already know this tournament approach. It's the core method ActiveGenie uses to deliver consistent results. I wrote the initial tournament logic back on February 3, 2025, nearly a year and a half ago.",{"type":9,"tag":10,"props":80,"children":81},{},[82],{"type":14,"value":83},"There is one key difference: ActiveGenie uses political debate while the paper uses logprobs. That topic deserves a dedicated post of its own.",{"type":9,"tag":10,"props":85,"children":86},{},[87],{"type":14,"value":88},"Until then, try ActiveGenie. Stop relying on expensive frontier models alone. Pair a leaner model with the right system architecture instead.",{"type":9,"tag":10,"props":90,"children":91},{},[92,94],{"type":14,"value":93},"Paper: ",{"type":9,"tag":95,"props":96,"children":100},"a",{"href":97,"rel":98},"https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.05391",[99],"nofollow",[101],{"type":14,"value":102},"LLM-as-a-Verifier",1788360944718]