Application + Bios — full text

Tap Copy on any block, paste, edit. Current wording, matched to the live sites.

Application · Main answer

I am an independent security researcher with an MS in computer science and a member of Anthropic's model safety bug bounty program on HackerOne. My work rests on one question: human review can't keep up with the internet's scale, so can several different models, reasoning together, ease that burden — handling far more scanning, fixing and patching, surfacing the high-signal findings — so a careful reviewer stays in the loop and covers more? My rule: another lab's model reviews each fix before it lands, I log the cross-family catches, and in my testing labs every break another lab finds becomes a test. The question this access supports: how much test-time compute, from which families and combinations, for models reasoning in concert to carry that load as reliably as a reviewer — not to replace them, but to raise how much they safely cover? Defense is the measurable start of the Alignment Hypothesis. I'll use it myself, on public open source and my own systems.

Application · Credentials & references

Anthropic model safety bug bounty program on HackerOne (associated with researcher Eric Buess; under NDA). Reference: Ado, who offered to refer me; a reference from Anthropic staff, and others, available on request.

Application · How I'll use the access

I will use Defense access myself from a dedicated account on public open-source code and my own systems: finding and validating vulnerabilities, generating patches, and running review rounds. My landing rule: a patch passes pre-run tests and two reviews by families that wrote neither the tests nor the fix. This access supports a matched-cost comparison of models working in concert against the strongest single family, measured by how much more a reviewer can safely cover. Adversarial red/blue rounds run as tests on my own code, and I've built a verifier for signed review receipts.

Anything found in public repositories goes privately to maintainers first. This work will test and refine the open protocol I plan to publish.

I understand Anthropic retains and monitors traffic under this access, without zero data retention, and that a CVP grant may not power client-facing services. Work on private code will run on separate permitted access, and I'll apply for productization approval.

Application · Background

Member of Anthropic's model safety bug bounty program on HackerOne (under NDA); a reference from Anthropic staff available on request. MS in computer science; frontier AI researcher since before ChatGPT. I keep concurrent subscriptions across major frontier labs for a multi-family build practice.

My rule is that a model from another family reviews a fix before it lands. In my practice, cross-family review catches what the authoring family misses: tests that can never fail, path-traversal holes, and builders reviewing their own work. A model unlike the one that wrote the code is the one most likely to break it. I keep a ledger of these catches.

In my own testing labs, a red/blue loop runs on my receipt kit: one lab's model attacks, another lab's model fixes, and every break becomes a test.

Direction: the Alignment Hypothesis, the question this work is built to test.

Application · Vision

The vision for si.build is an aligned-incentives movement and exchange where contributors point idle subscription compute at defending public open source, earning standing akin to service, without code access or merge rights. This individual Defense access is the first concrete step.

For private code, the runner encrypts code to an AWS Nitro Enclave and confirms each lab's terms; that enclave work is actively underway, with completion anticipated shortly.

The architecture is layered from the outside in: an isolated world with its own clock; a sandbox per job; a minimal gatekeeper granting narrow, expiring permissions; enclave-held keys; provenance captured outside the agents; and cross-family review rounds recorded as signed receipts, checked by the verifier I've built.

X bio

Trying to follow Jesus. MSCS. AI security: different models reasoning together to match a careful human reviewer at scale. Testing the Alignment Hypothesis.

LinkedIn headline

Independent AI-security researcher | Different models, reasoning together, to match a careful human reviewer at scale | Member, Anthropic model safety bug bounty (HackerOne) | Alignment Hypothesis | MSCS

LinkedIn headline (no bounty line)

Independent AI-security researcher | Different models, reasoning together, to match a careful human reviewer at scale | Alignment Hypothesis | MSCS

LinkedIn About

I work on AI security as an independent researcher, built on one question: can several genuinely different models, reasoning together, defend as reliably as a careful human reviewer at scale? Not copies of one mind, but models trained on different data, with different rewards and different weights. That difference is the point.

At asi.blue, my rule is that a model from another lab reviews a fix before it lands, so no family grades its own work. At asi.red, models attack my own code on purpose: one lab's model attacks, another lab's model fixes, and every break becomes a failing test and then a fix. A model unlike the one that wrote the code is the one most likely to break it. I keep a ledger of the cross-family catches.

The question at the heart of it, building on Together AI's Mixture of Agents and the research since: how much test-time compute, from which frontier and near-frontier model families, in what combinations, at which effort levels and in which harnesses, does it take for models working in concert to match or beat a careful human reviewer? And how do we keep improving that, so a good patch lands as fast as it's found, easing the review burden so people focus on the calls that most need them? I'm turning that into benchmarks, red team and blue team, from samples run in my own testing labs.

I work with frontier and near-frontier models from Claude, GPT, Gemini, Grok, DeepSeek and others. I work hands-free, by voice: I speak, and the system builds, reviews and ships.

I'm a member of Anthropic's model safety bug bounty program on HackerOne, and I hold a master's in computer science. I care deeply about AI safety.

All of it serves a larger question: my research, the Alignment Hypothesis, and a claim I call the Universal Alignment Imperative, that alignment may be a problem any civilization reaching for greater intelligence has to solve. It asks what a real training ground for values needs (genuine freedom, lasting consequences, and not knowing you're being tested) and how closely our world fits. si.build is the home for where this is going.

I'm going screenless myself, with an app I built for my own use that I can teach others to use. I write about AI on X (@EricBuess), give occasional talks for school districts and universities, and mentor high school students through the INSPIRE program at Grace Community School.

I know I hold some false beliefs. I just don't know which ones. I welcome conversations about AI safety, careful experiments and practical tools.

Short third-person bio

Eric Buess is an independent AI-security researcher and a member of Anthropic's model safety bug bounty program on HackerOne. His work rests on one question: can several genuinely different models, reasoning together, defend as reliably as a careful human reviewer at scale? At asi.blue, his rule is that a model from another lab reviews a fix before it lands. At asi.red, models from different labs attack his own code, and each break becomes a failing test and then a fix. He is turning one question into benchmarks: how much test-time compute, from which model families and in what combinations, it takes for models working in concert to match or beat a careful human reviewer. Underneath it all is his research, the Alignment Hypothesis. He holds a master's in computer science, works hands-free by voice, and is going screenless with an app he built for his own use.