A three-layer pipeline (rules-as-code, discrete choice model, agent-based model) for behavioural policy simulation that keeps what is measured separate from what is assumed, calibrated against a real adoption survey.
au-radar is a RADAR-consistent benchmark of how well AI describes and reaches Australian federal services. It scores Australia 5.32, and finds that a robots-respecting agent is blocked at the free legal database but not at the silent government register.
ClauseKit runs five bodies of law through an LLM extraction pipeline. 171 of 262 extracted rules never produce a yes or no, and that gap is the result.
Porting a fine-tuned sentence embedding model into a Drupal 10 module, running ONNX inference in-process via PHP FFI. No Ollama, no external API, no vector DB.
A decompose-then-verify pipeline that checks each claim in an AI output against its source document, and an honest look at what the RAGTruth numbers actually say about it.
How to use your existing Notion kanban board as the cockpit to manage your running AI agents, demonstrating human intervention and agent memory across runs.
Fine-tuned MiniLM exported to ONNX and run inside an existing Spring Boot app. No vector DB, no Bedrock per query, no GPU. The post I needed when I started.
How I built a two-mode spatial decision support tool that lets health planners ask plain-English questions about GP coverage gaps and get optimised facility placement recommendations with briefing-quality narrative.