A three-layer pipeline (rules-as-code, discrete choice model, agent-based model) for behavioural policy simulation that keeps what is measured separate from what is assumed, calibrated against a real adoption survey.
au-radar is a RADAR-consistent benchmark of how well AI describes and reaches Australian federal services. It scores Australia 5.32, and finds that a robots-respecting agent is blocked at the free legal database but not at the silent government register.
ClauseKit runs five bodies of law through an LLM extraction pipeline. 171 of 262 extracted rules never produce a yes or no, and that gap is the result.
A decompose-then-verify pipeline that checks each claim in an AI output against its source document, and an honest look at what the RAGTruth numbers actually say about it.
How to use your existing Notion kanban board as the cockpit to manage your running AI agents, demonstrating human intervention and agent memory across runs.
How I built a two-mode spatial decision support tool that lets health planners ask plain-English questions about GP coverage gaps and get optimised facility placement recommendations with briefing-quality narrative.