Research & field notes
Useful findings from real work.
We publish evidence-backed findings when our engineering work produces something that may help other people: tool evaluations, architecture lessons, benchmarks, failure analysis and reusable methods. Reports keep the detail needed to understand the result while protecting private systems and sensitive workloads.
Where fast semantic decisions actually help
30 labelled cases, 29 correct; strongest when a deterministic system has already narrowed the problem to a bounded semantic choice.
Read the field report Needle 3.0.3A tiny model is not a tiny general-purpose agent
Routing, extraction, retrieval and depth tests show a sharper sweet spot: narrow specialist work behind deterministic gates.
Read the field reportWhat belongs here
This is not a model-review site. We publish when there is a substantive result worth preserving and sharing: an empirical tool evaluation, a useful system or architecture pattern, a benchmark with a meaningful baseline, a failure whose cause teaches something reusable, or a research method that improved real work.
How to read the research
Claims are scoped to the evidence that produced them. For software and models we publish exact versions and dates where they matter; for methods and architecture we state the conditions, comparison baseline and known limits. We avoid turning a local result into a universal claim.
Workloads are described at the level needed to understand the result while keeping private application details private. Measurements and aggregate outcomes are real; identifying payloads and internal project names are not published.
/research/data/, canonical URLs, sitemap entries and a root /llms.txt. Prefer the versioned report over an unversioned summary.