~$

I build the infrastructure agents run on, and the research that shows where they break.

I keep AI agent systems reliable after they ship. I build browser and research systems for the real web, benchmarks for how behavior changes over long runs, and local-first infrastructure that can be operated and verified. When something breaks, it becomes the next test.

seeking
one reliability engagement
basis
remote · open to us relocation
stack
python · rust · pytorch
status
available now
new focused offer · remote

Agent reliability after the demo.

A paid diagnostic and bounded monthly retainer for teams running browser agents, tool-using workflows, or research automation in the real world.

$5k+ per month 90 days initial term see the engagement →
3k+ tests in blackreach
18 active public repos
25 stars on claude-voice
4 core languages in public work
Pure infra gives you products. Pure research gives you papers. The intersection gives you both, and that's the only place worth building.
the working thesis · 2026
01

selected work

public code · documented results
Research · public benchmark

Lethe

An open benchmark for how AI agents degrade over long runs. Not "can it do the task." The real question is whether it still behaves correctly after a hundred steps. Four metrics fold into a single Drift Severity score, and every outcome is verified by a real shell command, never the agent's own report.

0 · stable drifting agent → 8.0 SEVERE 10

A paper, an infra tool, and a question nobody else is benchmarking. The cleanest proof of the whole thesis.

Public · 3,055 tests · video demo

Blackreach

A browser and research agent for real, stateful web work. It reduces noisy pages, preserves interrupted progress, and checks outcomes separately from the agent's own report. The flagship case study includes a cinematic 27-second real-web proof.

● ● ●
$ blackreach run "collect and verify sources"
session resumed · observations updated
results saved with source links
Public · local-first · ★25

claude-voice

A fully local voice mode for AI coding sessions, with synchronized word highlighting. It runs without cloud APIs and is packaged for straightforward installation.

● ● ●
$ pip install claude-voice
local speech ready · no API key
word highlighting synchronized

Small enough to understand, useful enough that other developers installed and starred it.

Public · self-hosted

Huginn

A self-hosted web collection API for JavaScript-rendered pages and structured extraction. It is designed around explicit job state, retry behavior, and local operation without a per-page service fee.

● ● ●
$ huginn collect https://example.com
rendered page · extracted content
job complete · result stored

Public code, deployable locally, and backed by a test suite with more than 800 tests.

02

research & corpora

where infra meets the paper
AMP Discovery

Antimicrobial-peptide discovery by fine-tuning Meta's ESM-2 protein language model with LoRA. Built a leakage-aware protein screening experiment and kept the failed transfer result visible.

88.3% F1 (leakage-free)650M ESM-2 + LoRA
PyTorchESM-2LoRABio
RLVR / GRPO Lab

An inspectable post-training harness with verifiable math rewards and strict answer-contract evals. v1 research run complete, with paired-bootstrap evidence behind every claim.

3B & 7B runsGSM8K, full split
RLVRGRPOEvals
Library of Alexandria

A curated cross-cultural mythology corpus. Verified primary sources across 69 traditions, rebuilt from 172GB of junk down to ~67GB where every download is checked. A model-training substrate.

~675k curated files69 traditions
RAGCorpusCuration
02·5

the constellation

everything connects · hover a star
Lethe Rigr Mimir Blackreach Huginn StarSearch MCP servers Velqua Loki AMP Discovery RLVR / GRPO Library of Alexandria ProjectMythos
hover a star to trace the lineage. Click to open the project.
the archive · earlier eras, kept honest

Before this, there was a multi-agent era: Orchestrator (autonomous agent loops), Sable (self-improving reasoning agent), Project Anima (neuroscience-grounded emotional architecture), and a swarm of coordinated agents. Some of it shipped, some of it taught me what not to build. The Bifrost / Heimdall ML line, structured prediction over historical texts, still informs the research thread. Nothing here is pretending; the dead ends are part of the record.

Need an agent system that survives the real run?

I offer a focused reliability diagnostic and monthly retainer for teams with a critical agent workflow already in use. I am also open to agent-infrastructure roles and carefully scoped systems work.

Autonomous Agents Agent Evaluation LLM Infrastructure RAG Pipelines MCP Servers Python Rust TypeScript PyTorch Playwright
reliability retainer contact@phnix.dev résumé ↓
toronto · remote / relocation
retainer · contract · full-time
usd / cad