# Felipe Guedes (FGXDEV) > Felipe Guedes, Software Engineer and Systems Architect in Toledo and Cascavel, Paraná, Brazil. Full stack, System Design, audits and production AI. 600+ projects. I am Felipe Guedes. Software Engineer and Systems Architect, senior full stack in the literal sense: front end, back end and the infrastructure underneath, because a system is only as good as the layer nobody looked at. Systems Architect when the problem is scale, Principal Engineer for security when the problem is survival. I have taken part in more than 600 projects, building, reviewing or auditing them, from a neighbourhood print shop to a real-time game engine. The pattern never changes: the failure is there, in a boundary someone trusted, and I find it. It is not a talent. It is a childhood spent noticing the detail that decides whether you eat that week, turned into a method. I learn the way I learned at twelve, when a print shop owner showed me CorelDRAW once and two weeks later I was running the shop: watch, ask, try, fail, try again, own it. That is how electronics came, then product, code, cloud, security and the business side of running systems. Give me any company and, on the first conversation, I can tell you what its systems need and where they are going to break. Today I run FGXDEV. I am head of Cravamos, a news portal where one human and his AI agents publish with responsibility. I am building ELUCENIA, a global medical and scientific network to accelerate discovery, a mission that comes from my father, who spent forty years caring for patients. And I am building NextFoot, a football game, because a real-time simulation is the hardest system I know how to make simple. ## Location and services FGXDEV is the professional practice of Felipe Guedes, Software Engineer and Systems Architect based in Toledo and Cascavel, Paraná, Brazil: software development, architecture and System Design, software and security audits, and production AI for companies in western Paraná, São Paulo, Goiânia and other capitals, and remotely for clients in 19 countries. Lives in Toledo, Paraná, Brazil. On site in: Cascavel, Toledo, Foz do Iguaçu, Marechal Cândido Rondon, Medianeira, Palotina, Assis Chateaubriand, Umuarama, Francisco Beltrão, Pato Branco, Guarapuava. Remote in: Curitiba, Londrina, Maringá, Ponta Grossa, all of Brazil and abroad. Languages: Portuguese (Brazil) and English. ### Service pages (English / Portuguese) - [Software Engineer in Cascavel and Toledo, Paraná, Brazil](https://fgxdev.com/software-engineer-cascavel/) · PT-BR: https://fgxdev.com/pt/engenheiro-de-software-cascavel/ — Software Engineer and Systems Architect in Toledo and Cascavel, Paraná, Brazil. Full stack: front end, back end and infrastructure. 600+ projects. Working across Brazil and abroad. - [Custom software development in Cascavel and the region](https://fgxdev.com/software-development-cascavel/) · PT-BR: https://fgxdev.com/pt/desenvolvimento-de-software-cascavel/ — Custom software development in Cascavel, Toledo and western Paraná, Brazil: web systems, portals, APIs, multi-tenant platforms and integrations with Next.js, TypeScript, Node.js and PostgreSQL. - [Software architecture and System Design](https://fgxdev.com/software-architecture-system-design/) · PT-BR: https://fgxdev.com/pt/arquitetura-de-software-system-design/ — Systems Architect in Toledo and Cascavel, Paraná, Brazil. System Design for systems that must scale and fail safely: domain boundaries, data, queues, multi-tenancy, observability. Brazil and abroad. - [Software, architecture and security audit](https://fgxdev.com/software-audit-security/) · PT-BR: https://fgxdev.com/pt/auditoria-de-software-e-seguranca/ — Code, architecture and infrastructure audits from Toledo and Cascavel, Brazil, for clients anywhere. Finds the race condition, the data leak and the silent loss before they become incidents. Prioritised report. - [Production AI for companies](https://fgxdev.com/ai-for-companies/) · PT-BR: https://fgxdev.com/pt/inteligencia-artificial-para-empresas/ — Production AI for companies in Toledo, Cascavel, western Paraná and across Brazil: agents, RAG, automation with human review, evaluation and guardrails. Engineering, not demos. Eleven agents run Cravamos daily. - [Technology consulting in Toledo, Foz do Iguaçu and western Paraná](https://fgxdev.com/technology-consulting-west-parana/) · PT-BR: https://fgxdev.com/pt/consultoria-de-tecnologia-toledo-foz-do-iguacu/ — Technology and software engineering consulting for companies in Toledo, Foz do Iguaçu, Marechal Cândido Rondon, Medianeira, Umuarama and across western Paraná, Brazil. Diagnosis in the first conversation. - [Websites and systems for physicians, clinics and laboratories](https://fgxdev.com/websites-and-systems-for-clinics/) · PT-BR: https://fgxdev.com/pt/sites-e-sistemas-para-medicos-e-clinicas/ — Professional websites, scheduling, patient area and integrations for physicians, clinics and laboratories in Toledo, Cascavel and western Paraná, Brazil, with health-data security and privacy compliance. - [Technology and systems for industry in western Paraná](https://fgxdev.com/technology-for-industry-parana/) · PT-BR: https://fgxdev.com/pt/tecnologia-para-industria-oeste-do-parana/ — Systems, integrations and architecture for manufacturers in Toledo, Cascavel, Marechal Cândido Rondon, Medianeira and Palotina, Brazil: shop floor, ERP, real-time data, traceability and security. - [Systems for cooperatives and agribusiness in western Paraná](https://fgxdev.com/systems-for-cooperatives-and-agribusiness/) · PT-BR: https://fgxdev.com/pt/sistemas-para-cooperativas-e-agronegocio/ — Member portals, integrations, traceability and data for cooperatives and agribusiness companies in Toledo, Cascavel, Palotina, Marechal Cândido Rondon and western Paraná, Brazil. - [Professional websites for companies in Toledo and Cascavel](https://fgxdev.com/professional-websites-for-companies/) · PT-BR: https://fgxdev.com/pt/sites-profissionais-para-empresas-toledo-cascavel/ — Professional websites and portals for companies in Toledo, Cascavel and western Paraná, Brazil: fast, secure, with local SEO and ready for AI search. Built by a software engineer, not from a template. - All services: https://fgxdev.com/services/ · PT-BR: https://fgxdev.com/pt/services/ ## Questions and answers **Who is Felipe Guedes?** Felipe Guedes is a Software Engineer and Systems Architect specialised in System Design, based in Toledo and Cascavel, Paraná, Brazil. He has taken part in more than 600 projects, building, reviewing or auditing systems. He runs FGXDEV, is head of Cravamos and is building ELUCENIA and NextFoot. **Where does Felipe Guedes work?** In Toledo, where he lives, in Cascavel and western Paraná on site (Foz do Iguaçu, Marechal Cândido Rondon, Medianeira, Umuarama and the region), on site also in São Paulo, Goiânia and other capitals, and remotely across Brazil and abroad, in Portuguese and English. He has delivered projects for clients in 19 countries. **What does Felipe Guedes do?** Full-stack software development, software architecture and System Design, code, architecture, infrastructure and security audits, and production AI for companies. **How do I hire Felipe Guedes?** Through the contact form at fgxdev.com/contact/ or by email at contato@fgxdev.com. The first one-hour conversation is free and already ends with a diagnosis. **Who was Felipe Guedes' father?** Dr. Pedro Moretti Guedes (1935–2010), a physician in São Paulo for more than forty years who treated for free those who could not pay. Felipe maintains his memorial at pedromorettiguedes.com.br. **Do you work on site with companies in Cascavel?** Yes. I live in Toledo, western Paraná, 50 km from Cascavel, and hold in-person meetings in Cascavel, Toledo, Foz do Iguaçu and nearby cities, as well as São Paulo, Goiânia and other capitals. For the rest of Brazil and abroad the work is remote, with the same method and the same deliverables: projects in 19 countries so far. **What kind of company hires a Software Engineer and Systems Architect?** Companies that depend on a system to operate: cooperatives, industry, retail, healthcare, agribusiness, SaaS, portals and public bodies. Usually the system already exists and needs to scale, stop failing or be audited; sometimes it is about to be born and needs to be born right. **How does an engagement start?** With a one-hour conversation, free of charge, where I listen to the problem and tell you what your systems need and where they are going to break. After that, a proposal by fixed scope or by period, always with written deliverables. **Do you work with the team the company already has?** Yes, and I prefer it. Architecture and audits only matter if the team that maintains the system understands and agrees with the decisions. I document everything, review it with the team and leave the knowledge in house. **What is the difference between development, architecture and audit?** Development is building the system. Architecture is deciding how it is organised, scales and fails safely, before the first line of code. Audit is examining an existing system to find the failure nobody found. I do all three, separately or together. ## Highlights - System Design & architecture: How a system is organized, scales and fails safely, decided before the first line of code. - Full stack, end to end: Next.js, TypeScript, Node.js, PostgreSQL, and the Linux, Cloudflare and AWS layers under them. - Audit & security: Code, architecture and infrastructure reviews that find the race condition, the leak, the silent data loss. - AI in production: Agents, RAG, evals and guardrails shipped as systems, not demos. Eleven agents run Cravamos every day. - Business 360: Finance, legal, accounting and sales, trained at Biopark. I build companies, not just software. - Product & interface: From the data model to the pixel: interfaces that non-technical people operate alone, designed with the system underneath in mind. ## Cravamos One human. Eleven agents. Zero excuses. Cravamos is a news and services portal built the way the biggest technology companies in the world build: with an obsession for the reader, a standard for detail that does not negotiate, and systems that hold up when everyone is watching. We are not the size of Apple or Microsoft. We measure ourselves against them anyway. Not in headcount, in standard. The question we ask about every headline, every calculator, every pixel is the same one they ask: would we be proud if a hundred million people saw this today? Cravamos is one human and his AI agents. Eleven specialized agents research, draft, cross-check and format, around the clock, across sports, Formula 1, games, entertainment and everyday tools. One human, the head of Cravamos, holds the editorial responsibility for every word they publish. That arrangement is not a shortcut. It is a discipline. The agents make the scale possible. The human makes it accountable. Every piece carries identified sources, every tool respects the reader's data, and every mistake has a name attached to its correction. We refuse the click that costs the reader's trust. We refuse sources we cannot name. We refuse tools that harvest more than they help. We refuse to hide the machine behind the byline: how we use AI is written on the site, in plain language, for anyone to read. A portal that a whole country opens in the morning and trusts by lunch. Information, service and entertainment with human responsibility, at a scale no newsroom of one could reach before now, and at a quality no newsroom of any size should settle for. If we are not there yet, that is the direction. Everything we ship moves us one detail closer. ## Life story (summary) Son of a doctor I met only at fifteen. Raised by seven mothers. Working since I was nine. This is the whole story, no shortcuts. Chapters: A name I did not know was mine; Seven mothers; The cart; The print shop; Electronics; Two candidates and a test; Three months; Toledo; The bicycle; The agency; On foot; Biopark; Today. Full text: https://fgxdev.com/life/ ## Projects (current) - ELUCENIA [active] · Founder · Systems Architect · Scientific network · one mission: cancer: A global scientific network with one mission: cancer. Discovery Engine from signals to evidence, contradictions, gaps, hypotheses and investigations, every inference with its origin; 216 open source clinical tools; a public knowledge base. AI supports, science stays human. Stack: TypeScript, Next.js, Node.js, PostgreSQL, AI orchestration, Data pipelines, Provenance, Cloud. - Faultline [open source] · Creator · Maintainer · Open source · code auditor: Find the failure nobody finds. A zero-dependency auditor for the boundaries in a codebase: calls without timeouts, retries without backoff, money endpoints without idempotency keys, swallowed errors, SQL and shell built from strings, secrets in code, queries that forget the tenant. One command, a report with the fix for every finding, an exit code for CI and a GitHub Action. Stack: Node.js, Zero dependencies, CLI, GitHub Action, 22 rules, AGPL-3.0. - Facívia [building] · Systems Architect · Lead Engineer · B2B CRM platform: A multi-tenant CRM for B2B operations: pipeline, accounts, automation and integrations, designed for companies whose sales process is a system, not a spreadsheet. Stack: Next.js, TypeScript, Node.js, PostgreSQL, Multi-tenant, RBAC, REST & Webhooks, Observability. - FGX Web [active] · Principal Engineer · Security & Technology · Cloud, security and engineering: The engineering and security backbone behind the group's products: cloud architecture, hardening, audits, CI/CD, observability and incident response for systems that cannot go down. Stack: Cloudflare, AWS, Linux, CI/CD, Infrastructure as Code, Security audits, Observability, Incident response. - Klic.bio [active] · Founder · Architect · Web platform for entrepreneurs: Professional sites and software for entrepreneurs at a price that does not exclude anyone: multi-tenant platform, automated SEO, edge delivery and a publishing flow a non-technical owner can run alone. Stack: Next.js, TypeScript, Multi-tenant, Edge / CDN, SEO automation, Billing, PostgreSQL. - NextFoot [building] · Creator · Systems Architect · Football game · real-time simulation: A football game built from the engine up: match simulation, player and team modelling, real-time state and a multiplayer-ready architecture, where latency and consistency are decided by design, not by luck. Stack: TypeScript, Game engine, Real-time simulation, WebGL, Multiplayer netcode, Node.js, PostgreSQL. - Caullê [consulting] · Consultant · Technology, Development & Innovation · Digital product & innovation: Technology, development and innovation consulting: from product strategy and architecture to hands-on delivery of the systems that carry the business. Stack: Product strategy, Next.js, TypeScript, Cloud, AI, Architecture. - Cravamos [active] · Founder · Architect · AI-driven news portal: A news and services portal run by eleven specialized AI agents under human editorial responsibility: identified sources, daily coverage of sports, F1, games and entertainment, plus calculators and tools that respect user data. Stack: Next.js, React, TypeScript, 11 AI agents, Editorial pipeline, Cloudflare, Image optimization, SEO. - Sindilojas-GO [building] · Technology Lead · Trade association · members platform: Technology for the retail employers' union of Goiás, representing 56,000+ companies since 1945: a members portal with benefits club, partner area, market intelligence and secure authentication. Stack: Next.js, TypeScript, PHP, PostgreSQL, Authentication, Members area, Integrations, i18n. ## Track record 600+ projects: Built, reviewed or audited across ten years, in 19 countries. Retail, fintech, health, media, games, education, B2B SaaS, marketplaces and public-sector portals. Front end, back end and the infrastructure underneath. ## ELUCENIA Technology to accelerate scientific discovery. A global scientific network with one mission: cancer. Science moves at the speed of its slowest loop: form a hypothesis, gather evidence, test it, publish, repeat. Most of that loop is not thinking. It is searching, cleaning, reconciling and waiting. Those are the parts software is good at. ELUCENIA connects people, evidence and in-house agents in a scientific infrastructure that turns scattered knowledge into traceable, collaborative investigations open to human review. Papers, clinical trials, laboratory data, clinical observations and epidemiological signals describe parts of the same challenge; the scientific question connects them. The Discovery Engine structures the path from monitoring to scientific inquiry: signals, evidence, contradictions, gaps, hypotheses, investigations. Each inference keeps its origin, its limits and room for human evaluation. AI supports; science remains human. No promises of a cure or automatic discoveries. ## ELUCENIA open source tools (216 clinical calculators and scores, documentation in Portuguese) Free tools with zero dependencies and no data transfer, for any physician, clinic or agency to use on their own website. Each one reproduces a published formula or score, with sources, reference cases and tests, and is published in the ELUCENIA GitHub organization under the Apache-2.0 license. Documentation in Portuguese. Catalogue: https://fgxdev.com/elucenia/tools/ · PT-BR: https://fgxdev.com/pt/elucenia/ferramentas/ · Organization: https://github.com/Elucenia · License: Apache-2.0 · No independent clinical validation; they do not replace a physician's judgement. - Escore 4T (trombocitopenia induzida por heparina) (score; hematologia-e-hemoterapia, medicina-intensiva, cirurgia-cardiovascular): Probabilidade clínica de TIH. https://github.com/Elucenia/tool-4ts-trombocitopenia-por-heparina - Escore ABCD² (score; neurologia, medicina-de-emergencia, clinica-medica): Risco de AVC precoce após AIT. https://github.com/Elucenia/tool-abcd2 - ACR TI-RADS (calculator; radiologia-e-diagnostico-por-imagem, cirurgia-de-cabeca-e-pescoco, endocrinologia-e-metabologia): Nódulo de tireoide: categoria e indicação de PAAF. https://github.com/Elucenia/tool-acr-ti-rads - Alcoolemia estimada (fórmula de Widmark) (calculator; medicina-de-trafego, medicina-legal-e-pericia-medica, medicina-de-emergencia): Álcool no sangue estimado e limites do CTB. https://github.com/Elucenia/tool-alcoolemia-estimada - Índice de Aldrete modificado (score; anestesiologia, medicina-intensiva): Alta da recuperação pós-anestésica. https://github.com/Elucenia/tool-aldrete-modificado - Ânion gap (corrigido e delta-delta) (calculator; medicina-intensiva, medicina-de-emergencia, nefrologia): Ânion gap, correção pela albumina e razão delta. https://github.com/Elucenia/tool-anion-gap - APACHE II (calculator; medicina-intensiva, medicina-de-emergencia, anestesiologia): Gravidade e mortalidade prevista na UTI. https://github.com/Elucenia/tool-apache-ii - Escore de Apfel (NVPO) (score; anestesiologia, cirurgia-geral, cirurgia-plastica): Risco de náusea e vômito pós-operatórios. https://github.com/Elucenia/tool-apfel - Escore de Apgar (score; pediatria, ginecologia-e-obstetricia, anestesiologia): Avaliação do recém-nascido no 1º e 5º minutos. https://github.com/Elucenia/tool-apgar - APGAR familiar (score; medicina-de-familia-e-comunidade, geriatria, psiquiatria): Percepção da funcionalidade familiar (Smilkstein). https://github.com/Elucenia/tool-apgar-familiar - APRI (índice AST/plaquetas) (calculator; gastroenterologia, infectologia, clinica-medica): Fibrose significativa e cirrose na hepatite C. https://github.com/Elucenia/tool-apri - Área valvar aórtica (equação de continuidade) (calculator; cardiologia, cirurgia-cardiovascular): Estenose aórtica no ecocardiograma. https://github.com/Elucenia/tool-area-valvar-aortica - ARISCAT (score; cirurgia-toracica, anestesiologia, pneumologia): Risco de complicação pulmonar pós-operatória. https://github.com/Elucenia/tool-ariscat - ASRS v1.1 (rastreamento de TDAH no adulto) (calculator; psiquiatria, neurologia, medicina-de-familia-e-comunidade): As 6 perguntas de rastreamento de TDAH da OMS. https://github.com/Elucenia/tool-asrs-rastreamento-tdah - AUDIT (teste de identificação de problemas com álcool) (score; psiquiatria, medicina-de-familia-e-comunidade, clinica-medica): Uso de risco, nocivo e dependência de álcool. https://github.com/Elucenia/tool-audit - BED e EQD2 (calculator; radioterapia, oncologia-clinica): Dose biológica efetiva e equivalente em 2 Gy. https://github.com/Elucenia/tool-bed-e-eqd2 - Escore BISAP (score; gastroenterologia, cirurgia-do-aparelho-digestivo, medicina-de-emergencia): Gravidade da pancreatite aguda nas primeiras 24 h. https://github.com/Elucenia/tool-bisap - Boletim de Silverman-Andersen (score; pediatria, medicina-intensiva): Desconforto respiratório do recém-nascido. https://github.com/Elucenia/tool-boletim-de-silverman-andersen - Escala de Boston (BBPS) (calculator; endoscopia, coloproctologia, gastroenterologia): Qualidade do preparo intestinal na colonoscopia. https://github.com/Elucenia/tool-boston-bowel-preparation-scale - Questionário CAGE (score; psiquiatria, clinica-medica, medicina-de-familia-e-comunidade): Rastreamento rápido de abuso ou dependência de álcool. https://github.com/Elucenia/tool-cage - Cálcio corrigido pela albumina (calculator; medicina-de-emergencia, medicina-intensiva, endocrinologia-e-metabologia): Cálcio total ajustado à albumina sérica. https://github.com/Elucenia/tool-calcio-corrigido - Teste de caminhada de 6 minutos: distância prevista (calculator; pneumologia, cardiologia, medicina-fisica-e-reabilitacao): Previsto e limite inferior (Enright e Sherrill). https://github.com/Elucenia/tool-caminhada-de-6-minutos - Canadian C-Spine Rule (calculator; medicina-de-emergencia, ortopedia-e-traumatologia, neurocirurgia): Imagem cervical no trauma em paciente alerta. https://github.com/Elucenia/tool-canadian-c-spine-rule - Canadian CT Head Rule (calculator; medicina-de-emergencia, neurocirurgia, neurologia): Indicação de TC no trauma craniano leve. https://github.com/Elucenia/tool-canadian-ct-head-rule - Escore de Caprini (score; cirurgia-vascular, angiologia, cirurgia-geral): Risco de tromboembolismo venoso em cirurgia. https://github.com/Elucenia/tool-caprini - Carga tabágica (maços-ano) (calculator; pneumologia, medicina-de-familia-e-comunidade, cirurgia-toracica): Exposição ao cigarro e rastreamento de câncer de pulmão. https://github.com/Elucenia/tool-carga-tabagica - CDAI e SDAI (calculator; reumatologia, clinica-medica): Índices clínicos de atividade da artrite reumatoide. https://github.com/Elucenia/tool-cdai-sdai - Escore de Centor modificado (McIsaac) (score; infectologia, medicina-de-familia-e-comunidade, otorrinolaringologia): Probabilidade de faringite estreptocócica. https://github.com/Elucenia/tool-centor-mcisaac - CHA₂DS₂-VASc (calculator; cardiologia, neurologia, clinica-medica): Risco de AVC na fibrilação atrial. https://github.com/Elucenia/tool-cha2ds2-vasc - Child-Pugh (score; gastroenterologia, cirurgia-do-aparelho-digestivo, clinica-medica): Gravidade e prognóstico da cirrose. https://github.com/Elucenia/tool-child-pugh - CIWA-Ar (score; psiquiatria, medicina-de-emergencia, clinica-medica): Gravidade da síndrome de abstinência do álcool. https://github.com/Elucenia/tool-ciwa-ar - Classificação de Forrest (calculator; endoscopia, gastroenterologia, cirurgia-do-aparelho-digestivo): Úlcera péptica sangrante: ressangramento e conduta. https://github.com/Elucenia/tool-classificacao-de-forrest - Classificação de Tubiana (Dupuytren) (calculator; cirurgia-da-mao, ortopedia-e-traumatologia, cirurgia-plastica): Estádio da contratura de Dupuytren por raio. https://github.com/Elucenia/tool-classificacao-de-tubiana - Clearance de creatinina (Cockcroft-Gault) (calculator; nefrologia, clinica-medica, cardiologia): Ajuste de dose de medicamentos. https://github.com/Elucenia/tool-clearance-de-creatinina - Clearance de creatinina em urina de 24 h (calculator; nefrologia, clinica-medica, medicina-laboratorial): Função renal medida, com checagem da coleta. https://github.com/Elucenia/tool-clearance-de-creatinina-urina-24h - Contagem absoluta de neutrófilos (calculator; hematologia-e-hemoterapia, oncologia-clinica, medicina-laboratorial): Neutrófilos por µL e grau de neutropenia. https://github.com/Elucenia/tool-contagem-absoluta-de-neutrofilos - Controle dos sintomas da asma (GINA) (score; alergia-e-imunologia, pneumologia, pediatria): Asma bem controlada, parcialmente ou não controlada. https://github.com/Elucenia/tool-controle-da-asma-gina - Conversão de acuidade visual (calculator; oftalmologia, medicina-de-trafego, medicina-do-trabalho): Snellen, decimal e logMAR. https://github.com/Elucenia/tool-conversao-de-acuidade-visual - Conversão de unidades laboratoriais (calculator; medicina-laboratorial, clinica-medica, nefrologia): mg/dL ↔ mmol/L ou µmol/L (unidades SI). https://github.com/Elucenia/tool-conversao-de-unidades-laboratoriais - Correção do sódio (Adrogué-Madias) (calculator; medicina-intensiva, medicina-de-emergencia, nefrologia): Variação do sódio por litro de solução. https://github.com/Elucenia/tool-correcao-de-sodio-adrogue-madias - Critérios ACR/EULAR 2010 para artrite reumatoide (score; reumatologia, clinica-medica): Classificação de artrite reumatoide. https://github.com/Elucenia/tool-criterios-acr-eular-artrite-reumatoide - Critérios ACR/EULAR 2015 para gota (calculator; reumatologia, clinica-medica, ortopedia-e-traumatologia): Classificação de gota. https://github.com/Elucenia/tool-criterios-acr-eular-gota - Critérios de Duke-ISCVID 2023 (calculator; infectologia, cardiologia, clinica-medica): Diagnóstico de endocardite infecciosa. https://github.com/Elucenia/tool-criterios-de-duke - Critérios de Framingham para insuficiência cardíaca (calculator; cardiologia, clinica-medica, medicina-de-emergencia): Diagnóstico clínico de IC por critérios maiores e menores. https://github.com/Elucenia/tool-criterios-de-framingham-ic - Critérios de Light (calculator; pneumologia, cirurgia-toracica, clinica-medica): Derrame pleural: transudato ou exsudato. https://github.com/Elucenia/tool-criterios-de-light - Critérios de Ranson (score; cirurgia-do-aparelho-digestivo, gastroenterologia, cirurgia-geral): Prognóstico da pancreatite aguda na admissão e em 48 h. https://github.com/Elucenia/tool-criterios-de-ranson - CTS-6 (síndrome do túnel do carpo) (score; cirurgia-da-mao, ortopedia-e-traumatologia, neurologia): Probabilidade clínica de túnel do carpo. https://github.com/Elucenia/tool-cts-6 - CURB-65 (calculator; pneumologia, medicina-de-emergencia, infectologia): Gravidade da pneumonia comunitária (com CRB-65). https://github.com/Elucenia/tool-curb-65 - DAS28 (VHS e PCR) (calculator; reumatologia, clinica-medica): Atividade da artrite reumatoide em 28 articulações. https://github.com/Elucenia/tool-das28 - Decaimento radioativo (calculator; medicina-nuclear, radiologia-e-diagnostico-por-imagem, radioterapia): Atividade restante de um radiofármaco. https://github.com/Elucenia/tool-decaimento-radioativo - Déficit de água livre (calculator; medicina-intensiva, medicina-de-emergencia, nefrologia): Reposição de água na hipernatremia. https://github.com/Elucenia/tool-deficit-de-agua-livre - Déficit de ferro (Ganzoni) (calculator; hematologia-e-hemoterapia, clinica-medica, gastroenterologia): Dose total de ferro intravenoso. https://github.com/Elucenia/tool-deficit-de-ferro-ganzoni - Densidade do PSA (calculator; urologia, radiologia-e-diagnostico-por-imagem, oncologia-clinica): PSA ajustado ao volume da próstata. https://github.com/Elucenia/tool-densidade-do-psa - Dose de carboplatina (Calvert) (calculator; oncologia-clinica, ginecologia-e-obstetricia, pneumologia): Dose pela AUC-alvo e pela função renal. https://github.com/Elucenia/tool-dose-de-carboplatina - Dose de ruído (NR-15, Anexo 1) (calculator; medicina-do-trabalho, otorrinolaringologia): Soma C/T da exposição diária a ruído contínuo. https://github.com/Elucenia/tool-dose-de-ruido - Dose máxima de anestésico local (calculator; anestesiologia, medicina-de-emergencia, cirurgia-plastica): Lidocaína, bupivacaína e ropivacaína por peso. https://github.com/Elucenia/tool-dose-maxima-anestesico-local - Dose por superfície corporal (mg/m²) (calculator; oncologia-clinica, hematologia-e-hemoterapia, pediatria): Converte a dose em mg/m² na dose total em mg. https://github.com/Elucenia/tool-dose-por-superficie-corporal - Driving pressure e complacência estática (calculator; medicina-intensiva, anestesiologia, pneumologia): Pressão de distensão na ventilação mecânica. https://github.com/Elucenia/tool-driving-pressure - Infusão de drogas vasoativas (calculator; medicina-intensiva, medicina-de-emergencia, anestesiologia): mcg/kg/min ↔ mL/h pela concentração. https://github.com/Elucenia/tool-drogas-vasoativas - EASI (Eczema Area and Severity Index) (calculator; dermatologia, alergia-e-imunologia, pediatria): Sinais e extensão da dermatite atópica. https://github.com/Elucenia/tool-easi - ECOG e Karnofsky (calculator; oncologia-clinica, radioterapia, hematologia-e-hemoterapia): Converte o Karnofsky em ECOG (performance status). https://github.com/Elucenia/tool-ecog-karnofsky - Equivalência de corticoides (calculator; clinica-medica, reumatologia, endocrinologia-e-metabologia): Conversão de doses entre glicocorticoides sistêmicos. https://github.com/Elucenia/tool-equivalencia-de-corticoides - Equivalente esférico e transposição (calculator; oftalmologia, pediatria): Refração em EE, transposição e vetores J0/J45. https://github.com/Elucenia/tool-equivalente-esferico - Escala Clínica de Fragilidade (CFS) (score; geriatria, medicina-intensiva, medicina-de-emergencia): Grau de fragilidade da pessoa idosa, de 1 a 9. https://github.com/Elucenia/tool-escala-clinica-de-fragilidade - Escala de Ashworth modificada (calculator; medicina-fisica-e-reabilitacao, neurologia, ortopedia-e-traumatologia): Graduação clínica da espasticidade. https://github.com/Elucenia/tool-escala-de-ashworth-modificada - Escala de risco familiar de Coelho-Savassi (calculator; medicina-de-familia-e-comunidade, medicina-preventiva-e-social): Risco familiar para priorizar visitas domiciliares. https://github.com/Elucenia/tool-escala-de-coelho-savassi - Escala de Coma de Glasgow (com GCS-P) (calculator; medicina-de-emergencia, medicina-intensiva, neurologia): Nível de consciência e reatividade pupilar. https://github.com/Elucenia/tool-escala-de-coma-de-glasgow - Escala de Depressão Pós-Parto de Edimburgo (EPDS) (calculator; ginecologia-e-obstetricia, psiquiatria, medicina-de-familia-e-comunidade): Rastreamento de depressão na gestação e no pós-parto. https://github.com/Elucenia/tool-escala-de-edimburgo-epds - Escala de Fisher modificada (score; neurocirurgia, neurologia, radiologia-e-diagnostico-por-imagem): Risco de vasoespasmo na hemorragia subaracnóidea. https://github.com/Elucenia/tool-escala-de-fisher-modificada - Escala de Hoehn e Yahr (score; neurologia, geriatria, medicina-fisica-e-reabilitacao): Estadiamento da doença de Parkinson. https://github.com/Elucenia/tool-escala-de-hoehn-e-yahr - Escala de House-Brackmann (calculator; otorrinolaringologia, cirurgia-de-cabeca-e-pescoco, neurologia): Grau de disfunção do nervo facial. https://github.com/Elucenia/tool-escala-de-house-brackmann - Escala de Hunt-Hess (score; neurocirurgia, neurologia, medicina-intensiva): Gravidade clínica da hemorragia subaracnóidea. https://github.com/Elucenia/tool-escala-de-hunt-hess - Índice de Katz (ABVD) (score; geriatria, medicina-fisica-e-reabilitacao, medicina-de-familia-e-comunidade): Independência nas atividades básicas de vida diária. https://github.com/Elucenia/tool-escala-de-katz - Escala de Lawton-Brody (AIVD) (calculator; geriatria, medicina-fisica-e-reabilitacao, medicina-de-familia-e-comunidade): Atividades instrumentais de vida diária. https://github.com/Elucenia/tool-escala-de-lawton - Escala de Rankin modificada (mRS) (score; neurologia, medicina-fisica-e-reabilitacao, neurocirurgia): Grau de incapacidade após o AVC. https://github.com/Elucenia/tool-escala-de-rankin-modificada - Escala de Resultados de Glasgow (GOS) (score; neurocirurgia, medicina-intensiva, medicina-fisica-e-reabilitacao): Desfecho funcional após lesão cerebral. https://github.com/Elucenia/tool-escala-de-resultados-de-glasgow - Escala de dor LANSS (score; neurologia, anestesiologia, acupuntura): Rastreamento de dor neuropática. https://github.com/Elucenia/tool-escala-lanss - Escala WFNS (calculator; neurocirurgia, neurologia, medicina-intensiva): Grau da hemorragia subaracnóidea pelo Glasgow. https://github.com/Elucenia/tool-escala-wfns - Escore AIR (Appendicitis Inflammatory Response) (score; cirurgia-geral, medicina-de-emergencia, cirurgia-pediatrica): Apendicite aguda com leucócitos, neutrófilos e PCR. https://github.com/Elucenia/tool-escore-air-apendicite - Escore de Alvarado (score; cirurgia-geral, medicina-de-emergencia, cirurgia-pediatrica): Probabilidade de apendicite aguda. https://github.com/Elucenia/tool-escore-de-alvarado - Escore de Beighton (calculator; ortopedia-e-traumatologia, reumatologia, medicina-esportiva): Hipermobilidade articular generalizada. https://github.com/Elucenia/tool-escore-de-beighton - Escore de Duke (esteira) (calculator; cardiologia, medicina-esportiva): Prognóstico no teste ergométrico. https://github.com/Elucenia/tool-escore-de-duke - Escore de Genebra revisado (calculator; pneumologia, medicina-de-emergencia, cardiologia): Probabilidade clínica de TEP (com a versão simplificada). https://github.com/Elucenia/tool-escore-de-genebra - Escore de Mayo parcial (retocolite) (score; coloproctologia, gastroenterologia): Atividade clínica da retocolite ulcerativa sem endoscopia. https://github.com/Elucenia/tool-escore-de-mayo-parcial - Escore de Mirels (score; ortopedia-e-traumatologia, oncologia-clinica, radioterapia): Risco de fratura patológica em metástase óssea. https://github.com/Elucenia/tool-escore-de-mirels - Escore de Risco Global (calculator; cardiologia, medicina-de-familia-e-comunidade, clinica-medica): Risco de eventos cardiovasculares em 10 anos. https://github.com/Elucenia/tool-escore-de-risco-global - Escore de Rockall (calculator; endoscopia, gastroenterologia, medicina-de-emergencia): Risco de óbito na hemorragia digestiva alta. https://github.com/Elucenia/tool-escore-de-rockall - Escore de Villalta (calculator; angiologia, cirurgia-vascular, hematologia-e-hemoterapia): Diagnóstico e gravidade da síndrome pós-trombótica. https://github.com/Elucenia/tool-escore-de-villalta - Escore de Westley (crupe) (score; pediatria, medicina-de-emergencia, otorrinolaringologia): Gravidade da laringotraqueobronquite. https://github.com/Elucenia/tool-escore-de-westley - Escore de Wexner (incontinência fecal) (score; coloproctologia, cirurgia-geral, ginecologia-e-obstetricia): Gravidade da incontinência anal (Cleveland Clinic). https://github.com/Elucenia/tool-escore-de-wexner - Escore isquêmico de Hachinski (score; neurologia, geriatria, psiquiatria): Demência vascular ou degenerativa. https://github.com/Elucenia/tool-escore-isquemico-de-hachinski - MACIS (carcinoma papilífero de tireoide) (calculator; cirurgia-de-cabeca-e-pescoco, endocrinologia-e-metabologia, cirurgia-oncologica): Prognóstico do carcinoma papilífero de tireoide. https://github.com/Elucenia/tool-escore-macis - MESS (Mangled Extremity Severity Score) (calculator; ortopedia-e-traumatologia, cirurgia-vascular, cirurgia-geral): Gravidade do membro inferior esmagado. https://github.com/Elucenia/tool-escore-mess - Escore de força muscular MRC (soma) (score; medicina-fisica-e-reabilitacao, medicina-intensiva, neurologia): Força global e fraqueza adquirida na UTI. https://github.com/Elucenia/tool-escore-mrc - Escore TWIST (torção testicular) (score; cirurgia-pediatrica, urologia, medicina-de-emergencia): Risco de torção testicular no escroto agudo. https://github.com/Elucenia/tool-escore-twist - Estadiamento KDIGO da doença renal crônica (calculator; nefrologia, clinica-medica, medicina-de-familia-e-comunidade): Categoria G/A e risco pelo mapa de calor. https://github.com/Elucenia/tool-estadiamento-kdigo - Estatura por ossos longos (Trotter e Gleser) (calculator; medicina-legal-e-pericia-medica, patologia): Estimativa da estatura a partir do comprimento ósseo. https://github.com/Elucenia/tool-estatura-por-ossos-longos - Etilômetro: valor considerado (Resolução Contran 432) (calculator; medicina-de-trafego, medicina-legal-e-pericia-medica): Medição do bafômetro, erro admissível e enquadramento no CTB. https://github.com/Elucenia/tool-etilometro-ctb - EuroSCORE II (calculator; cirurgia-cardiovascular, cardiologia, anestesiologia): Mortalidade prevista em cirurgia cardíaca. https://github.com/Elucenia/tool-euroscore-ii - FIB-4 (fibrose hepática) (calculator; gastroenterologia, clinica-medica, infectologia): Fibrose hepática avançada sem biópsia. https://github.com/Elucenia/tool-fib-4 - FINDRISC (score; endocrinologia-e-metabologia, medicina-de-familia-e-comunidade, clinica-medica): Risco de diabetes tipo 2 em 10 anos. https://github.com/Elucenia/tool-findrisc - Fleischner 2017 (nódulo pulmonar sólido) (calculator; radiologia-e-diagnostico-por-imagem, pneumologia, cirurgia-toracica): Seguimento do nódulo pulmonar incidental na TC. https://github.com/Elucenia/tool-fleischner-2017 - Fórmula de Parkland (calculator; cirurgia-plastica, medicina-de-emergencia, medicina-intensiva): Reposição volêmica no grande queimado. https://github.com/Elucenia/tool-formula-de-parkland - Fração de ejeção (Teichholz) e encurtamento (calculator; cardiologia, radiologia-e-diagnostico-por-imagem): A partir dos diâmetros do VE. https://github.com/Elucenia/tool-fracao-de-ejecao-teichholz - Fração de excreção de sódio (FENa) (calculator; nefrologia, clinica-medica, medicina-intensiva): Pré-renal ou necrose tubular aguda na LRA oligúrica. https://github.com/Elucenia/tool-fracao-de-excrecao-de-sodio - Fração de excreção de ureia (FEUr) (calculator; nefrologia, clinica-medica, medicina-intensiva): LRA pré-renal ou NTA em quem usou diurético. https://github.com/Elucenia/tool-fracao-de-excrecao-de-ureia - FC máxima prevista e índice cronotrópico (calculator; medicina-esportiva, cardiologia): Resposta cronotrópica no esforço. https://github.com/Elucenia/tool-frequencia-cardiaca-maxima - Função discriminante de Maddrey (calculator; gastroenterologia, clinica-medica, medicina-intensiva): Gravidade da hepatite alcoólica e corticoide. https://github.com/Elucenia/tool-funcao-discriminante-de-maddrey - Escala GAD-7 (score; psiquiatria, medicina-de-familia-e-comunidade, clinica-medica): Rastreamento e gravidade da ansiedade. https://github.com/Elucenia/tool-gad-7 - Ganho de peso gestacional (IOM 2009) (calculator; ginecologia-e-obstetricia, nutrologia, medicina-de-familia-e-comunidade): Faixa recomendada pelo IMC pré-gestacional. https://github.com/Elucenia/tool-ganho-de-peso-gestacional - Interpretação da gasometria arterial (calculator; medicina-intensiva, medicina-de-emergencia, nefrologia): Distúrbio primário e compensação esperada. https://github.com/Elucenia/tool-gasometria-arterial - Gasto energético (Mifflin-St Jeor e Harris-Benedict) (calculator; nutrologia, endocrinologia-e-metabologia, clinica-medica): Gasto de repouso e gasto total com nível de atividade. https://github.com/Elucenia/tool-gasto-energetico - Gasto energético por METs (calculator; medicina-esportiva, nutrologia, endocrinologia-e-metabologia): Calorias e MET-minutos de uma atividade. https://github.com/Elucenia/tool-gasto-energetico-por-mets - Escala de Depressão Geriátrica (GDS-15) (score; geriatria, psiquiatria, medicina-de-familia-e-comunidade): Rastreamento de depressão na pessoa idosa. https://github.com/Elucenia/tool-gds-15 - Escore de Glasgow-Blatchford (score; endoscopia, gastroenterologia, medicina-de-emergencia): Hemorragia digestiva alta: alta precoce ou intervenção. https://github.com/Elucenia/tool-glasgow-blatchford - Gleason e grupo de grau ISUP (calculator; patologia, urologia, oncologia-clinica): Converte o escore de Gleason no grupo de grau (1 a 5). https://github.com/Elucenia/tool-gleason-isup - Classificação GOLD da DPOC (calculator; pneumologia, clinica-medica, medicina-de-familia-e-comunidade): Gravidade espirométrica (GOLD 1–4) e grupo A, B ou E. https://github.com/Elucenia/tool-gold-dpoc - Escore GRACE (mortalidade hospitalar) (calculator; cardiologia, medicina-de-emergencia, medicina-intensiva): Risco de morte na síndrome coronariana aguda. https://github.com/Elucenia/tool-grace - Gradiente alvéolo-arterial de O₂ (calculator; medicina-intensiva, medicina-de-emergencia, pneumologia): Causa da hipoxemia: pulmão ou hipoventilação. https://github.com/Elucenia/tool-gradiente-alveolo-arterial - Média tonal e grau de perda auditiva (OMS 2021) (calculator; otorrinolaringologia, medicina-do-trabalho, geriatria): Médias tritonal e quadritonal e grau da OMS. https://github.com/Elucenia/tool-grau-de-perda-auditiva - Gravidade da anafilaxia (Brown) (calculator; alergia-e-imunologia, medicina-de-emergencia, pediatria): Classificação em leve, moderada ou grave. https://github.com/Elucenia/tool-gravidade-da-anafilaxia - Escore H₂FPEF (score; cardiologia, clinica-medica): Probabilidade de IC com fração de ejeção preservada. https://github.com/Elucenia/tool-h2fpef - Equilíbrio de Hardy-Weinberg (calculator; genetica-medica, ginecologia-e-obstetricia, pediatria): Portadores e risco do casal em doença recessiva. https://github.com/Elucenia/tool-hardy-weinberg - Índice de Harvey-Bradshaw (calculator; gastroenterologia, coloproctologia): Atividade clínica da doença de Crohn. https://github.com/Elucenia/tool-harvey-bradshaw - HAS-BLED (calculator; cardiologia, hematologia-e-hemoterapia, clinica-medica): Risco de sangramento com anticoagulação. https://github.com/Elucenia/tool-has-bled - HbA1c e glicemia média estimada (ADAG) (calculator; endocrinologia-e-metabologia, clinica-medica, medicina-de-familia-e-comunidade): Conversão entre hemoglobina glicada e glicemia média. https://github.com/Elucenia/tool-hba1c-glicemia-media-estimada - Escore HEART (calculator; cardiologia, medicina-de-emergencia): Dor torácica na emergência. https://github.com/Elucenia/tool-heart-score - Hidratação de manutenção (Holliday-Segar) (calculator; pediatria, cirurgia-pediatrica, medicina-intensiva): Volume diário e regra 4-2-1. https://github.com/Elucenia/tool-holliday-segar - HOMA-IR e HOMA-β (calculator; endocrinologia-e-metabologia, nutrologia, clinica-medica): Resistência à insulina e função da célula beta. https://github.com/Elucenia/tool-homa-ir - Hipertrofia ventricular esquerda no ECG (calculator; cardiologia, clinica-medica, medicina-do-trabalho): Sokolow-Lyon, Cornell voltagem e produto de Cornell. https://github.com/Elucenia/tool-hve-no-ecg - IBUTG e limite de exposição ao calor (NR-15) (calculator; medicina-do-trabalho): Índice de bulbo úmido termômetro de globo. https://github.com/Elucenia/tool-ibutg - ICH Score (score; neurologia, neurocirurgia, medicina-intensiva): Mortalidade em 30 dias na hemorragia intracerebral. https://github.com/Elucenia/tool-ich-score - Idade corrigida do prematuro (calculator; pediatria, medicina-de-familia-e-comunidade): Idade corrigida e idade pós-menstrual. https://github.com/Elucenia/tool-idade-corrigida-do-prematuro - Idade gestacional e DPP pela DUM (calculator; ginecologia-e-obstetricia, medicina-de-familia-e-comunidade): Regra de Naegele e idade gestacional em semanas. https://github.com/Elucenia/tool-idade-gestacional-e-dpp - Idade gestacional pelo CCN (calculator; ginecologia-e-obstetricia, radiologia-e-diagnostico-por-imagem): Datação pelo comprimento cabeça-nádega (Robinson). https://github.com/Elucenia/tool-idade-gestacional-pelo-ccn - IMC (Índice de Massa Corporal) (calculator; clinica-medica, nutrologia, endocrinologia-e-metabologia): Classificação da OMS e pontos de corte para asiáticos. https://github.com/Elucenia/tool-imc - Índice BODE (calculator; pneumologia, cirurgia-toracica, medicina-fisica-e-reabilitacao): Prognóstico multidimensional na DPOC. https://github.com/Elucenia/tool-indice-bode - Índice de Barthel (score; medicina-fisica-e-reabilitacao, geriatria, neurologia): Independência nas atividades básicas de vida diária. https://github.com/Elucenia/tool-indice-de-barthel - Índice de Baux revisado (calculator; cirurgia-plastica, medicina-intensiva, medicina-de-emergencia): Gravidade e prognóstico no grande queimado. https://github.com/Elucenia/tool-indice-de-baux-revisado - Índice de Bishop (score; ginecologia-e-obstetricia): Colo uterino favorável para indução do parto. https://github.com/Elucenia/tool-indice-de-bishop - Índice de Capacidade para o Trabalho (ICT) (calculator; medicina-do-trabalho, medicina-fisica-e-reabilitacao): Work Ability Index, versão brasileira. https://github.com/Elucenia/tool-indice-de-capacidade-para-o-trabalho - Índice de Comorbidade de Charlson (calculator; clinica-medica, geriatria, oncologia-clinica): Carga de comorbidades e sobrevida em 10 anos. https://github.com/Elucenia/tool-indice-de-charlson - Índice de choque (e modificado) (calculator; medicina-de-emergencia, medicina-intensiva, cirurgia-geral): FC/PAS e FC/PAM no trauma e no choque. https://github.com/Elucenia/tool-indice-de-choque - Índice de Mentzer (calculator; hematologia-e-hemoterapia, medicina-laboratorial, pediatria): Traço talassêmico ou anemia ferropriva?. https://github.com/Elucenia/tool-indice-de-mentzer - Reticulócitos corrigidos e IPR (calculator; hematologia-e-hemoterapia, medicina-laboratorial, clinica-medica): Resposta medular na anemia. https://github.com/Elucenia/tool-indice-de-producao-reticulocitaria - Índice de Risco Cardíaco Revisado (Lee) (score; anestesiologia, cardiologia, clinica-medica): Risco cardíaco em cirurgia não cardíaca. https://github.com/Elucenia/tool-indice-de-risco-cardiaco-revisado - Índice de Tobin (respiração rápida e superficial) (calculator; medicina-intensiva, pneumologia, anestesiologia): Preditor de sucesso no desmame ventilatório. https://github.com/Elucenia/tool-indice-de-tobin - Índice Prognóstico de Van Nuys (USC/VNPI) (score; mastologia, radioterapia, cirurgia-oncologica): Recidiva local do carcinoma ductal in situ. https://github.com/Elucenia/tool-indice-de-van-nuys - Índice Prognóstico de Nottingham (NPI) (calculator; mastologia, oncologia-clinica, cirurgia-oncologica): Prognóstico do câncer de mama operável. https://github.com/Elucenia/tool-indice-prognostico-de-nottingham - Índice tornozelo-braquial (ITB) (calculator; angiologia, cirurgia-vascular, cardiologia): Doença arterial periférica. https://github.com/Elucenia/tool-indice-tornozelo-braquial - Intervalo post-mortem pela temperatura (Henssge) (calculator; medicina-legal-e-pericia-medica, patologia): Tempo de morte pela temperatura retal. https://github.com/Elucenia/tool-intervalo-post-mortem-henssge - IPI (Índice Prognóstico Internacional) (score; hematologia-e-hemoterapia, oncologia-clinica): Prognóstico do linfoma não Hodgkin agressivo. https://github.com/Elucenia/tool-ipi-linfoma - IPSS (Escore Internacional de Sintomas Prostáticos) (score; urologia, geriatria, medicina-de-familia-e-comunidade): Gravidade dos sintomas urinários na HPB. https://github.com/Elucenia/tool-ipss - Escore ISTH de CID manifesta (score; hematologia-e-hemoterapia, medicina-intensiva, medicina-laboratorial): Coagulação intravascular disseminada (ISTH). https://github.com/Elucenia/tool-isth-cid - KFRE (Kidney Failure Risk Equation) (calculator; nefrologia, clinica-medica, medicina-de-familia-e-comunidade): Risco de falência renal em 2 e 5 anos. https://github.com/Elucenia/tool-kfre - Escore de Khorana (score; oncologia-clinica, hematologia-e-hemoterapia, cirurgia-oncologica): Risco de TEV no paciente em quimioterapia. https://github.com/Elucenia/tool-khorana - Kt/V e URR na hemodiálise (calculator; nefrologia, medicina-intensiva): Adequação da dose de hemodiálise (Daugirdas). https://github.com/Elucenia/tool-kt-v-hemodialise - LDL-colesterol calculado (calculator; cardiologia, endocrinologia-e-metabologia, medicina-laboratorial): Friedewald, Sampson (NIH) e não-HDL. https://github.com/Elucenia/tool-ldl-calculado - Fórmula SRK II (lente intraocular) (calculator; oftalmologia): Poder da LIO pela fórmula SRK II (didático). https://github.com/Elucenia/tool-lente-intraocular-srk-ii - Índice MASCC (score; oncologia-clinica, hematologia-e-hemoterapia, infectologia): Baixo risco na neutropenia febril. https://github.com/Elucenia/tool-mascc - Massa do VE e geometria ventricular (calculator; cardiologia, radiologia-e-diagnostico-por-imagem): Fórmula da ASE (Devereux), índice e ERP. https://github.com/Elucenia/tool-massa-ventricular-esquerda - MELD-Na e MELD 3.0 (calculator; gastroenterologia, cirurgia-do-aparelho-digestivo, clinica-medica): Mortalidade em 90 dias na cirrose e fila de transplante. https://github.com/Elucenia/tool-meld - Método de Capurro (somático) (calculator; pediatria, ginecologia-e-obstetricia): Idade gestacional do recém-nascido. https://github.com/Elucenia/tool-metodo-de-capurro - NAFLD Fibrosis Score (NFS) (calculator; gastroenterologia, endocrinologia-e-metabologia, clinica-medica): Fibrose avançada na esteatose hepática. https://github.com/Elucenia/tool-nafld-fibrosis-score - NEWS2 (National Early Warning Score 2) (calculator; medicina-de-emergencia, medicina-intensiva, clinica-medica): Deterioração clínica pelos sinais vitais. https://github.com/Elucenia/tool-news2 - Critérios NEXUS (coluna cervical) (score; medicina-de-emergencia, ortopedia-e-traumatologia, neurocirurgia): Dispensa de imagem cervical no trauma fechado. https://github.com/Elucenia/tool-nexus-coluna-cervical - NIHSS (escala de AVC do NIH) (score; neurologia, medicina-de-emergencia, medicina-intensiva): Gravidade do déficit neurológico no AVC agudo. https://github.com/Elucenia/tool-nihss - Grau histológico de Nottingham (score; patologia, mastologia, oncologia-clinica): Grau do carcinoma invasivo de mama (Elston-Ellis). https://github.com/Elucenia/tool-nottingham - NRS-2002 (score; nutrologia, clinica-medica, medicina-intensiva): Triagem de risco nutricional no paciente internado. https://github.com/Elucenia/tool-nrs-2002 - NNT e NNH (número necessário para tratar) (calculator; medicina-preventiva-e-social, medicina-de-familia-e-comunidade, clinica-medica): NNT, RRA, RRR e risco relativo de um ensaio clínico. https://github.com/Elucenia/tool-numero-necessario-para-tratar - Osmolaridade sérica e gap osmolar (calculator; medicina-de-emergencia, medicina-intensiva, nefrologia): Osmolaridade calculada, efetiva e gap osmolar. https://github.com/Elucenia/tool-osmolaridade-serica - Escore de Pádua (score; hematologia-e-hemoterapia, clinica-medica, oncologia-clinica): Risco de TEV no paciente clínico internado. https://github.com/Elucenia/tool-padua - PASI (Psoriasis Area and Severity Index) (calculator; dermatologia, reumatologia): Gravidade da psoríase e resposta ao tratamento. https://github.com/Elucenia/tool-pasi - Pediatric Appendicitis Score (PAS) (score; cirurgia-pediatrica, pediatria, medicina-de-emergencia): Probabilidade de apendicite na criança. https://github.com/Elucenia/tool-pediatric-appendicitis-score - Critérios PERC (score; medicina-de-emergencia, pneumologia, clinica-medica): Exclusão clínica de TEP sem D-dímero. https://github.com/Elucenia/tool-perc - Peso fetal estimado (Hadlock) (calculator; ginecologia-e-obstetricia, radiologia-e-diagnostico-por-imagem): Peso fetal pela biometria e percentil para a idade. https://github.com/Elucenia/tool-peso-fetal-estimado-hadlock - Peso ideal e peso ajustado (calculator; nutrologia, clinica-medica, medicina-intensiva): Fórmula de Devine e peso ajustado para obesos. https://github.com/Elucenia/tool-peso-ideal-e-ajustado - Peso predito e volume corrente protetor (calculator; medicina-intensiva, anestesiologia, medicina-de-emergencia): Volume corrente de 6 mL/kg (ARDSNet). https://github.com/Elucenia/tool-peso-predito-volume-corrente - Questionário PHQ-9 (calculator; psiquiatria, medicina-de-familia-e-comunidade, clinica-medica): Rastreamento e gravidade da depressão. https://github.com/Elucenia/tool-phq-9 - Pressão arterial média e de pulso (calculator; cardiologia, medicina-intensiva, clinica-medica): PAM, pressão de pulso e estágio da HA. https://github.com/Elucenia/tool-pressao-arterial-media - Probabilidade pós-teste (teorema de Bayes) (calculator; medicina-preventiva-e-social, medicina-de-familia-e-comunidade, clinica-medica): Da probabilidade pré-teste à pós-teste pela razão de verossimilhança. https://github.com/Elucenia/tool-probabilidade-pos-teste - PSI/PORT (índice de gravidade da pneumonia) (calculator; pneumologia, infectologia, medicina-de-emergencia): Classe de risco de Fine na pneumonia comunitária. https://github.com/Elucenia/tool-psi-port - qSOFA (quick SOFA) (score; medicina-de-emergencia, medicina-intensiva, infectologia): Rastreio de gravidade na suspeita de sepse. https://github.com/Elucenia/tool-qsofa - QT corrigido (QTc) (calculator; cardiologia, medicina-de-emergencia, psiquiatria): Bazett, Fridericia, Framingham e Hodges. https://github.com/Elucenia/tool-qtc - Escala de Agitação e Sedação de Richmond (RASS) (score; medicina-intensiva, anestesiologia, medicina-de-emergencia): Nível de sedação e agitação na UTI. https://github.com/Elucenia/tool-rass - Regra de Ottawa para joelho (score; ortopedia-e-traumatologia, medicina-de-emergencia, medicina-esportiva): Quando pedir radiografia no trauma do joelho. https://github.com/Elucenia/tool-regra-de-ottawa-joelho - Regras de Ottawa para tornozelo e pé (calculator; ortopedia-e-traumatologia, medicina-de-emergencia, medicina-esportiva): Quando pedir radiografia no entorse de tornozelo. https://github.com/Elucenia/tool-regras-de-ottawa-tornozelo - Relação PaO₂/FiO₂ e SDRA (Berlim) (calculator; medicina-intensiva, medicina-de-emergencia, pneumologia): Oxigenação e gravidade da SDRA. https://github.com/Elucenia/tool-relacao-pao2-fio2 - 1RM estimada (Epley e Brzycki) (calculator; medicina-esportiva, medicina-fisica-e-reabilitacao): Uma repetição máxima sem teste máximo. https://github.com/Elucenia/tool-repeticao-maxima-1rm - Risco de síndrome de Down pela idade materna (calculator; genetica-medica, ginecologia-e-obstetricia): Risco basal de trissomia 21 ao nascimento. https://github.com/Elucenia/tool-risco-de-trissomia-21-pela-idade-materna - Risco relativo e odds ratio (calculator; medicina-preventiva-e-social, medicina-de-familia-e-comunidade, medicina-do-trabalho): Medidas de associação da tabela 2×2 com IC 95%. https://github.com/Elucenia/tool-risco-relativo-e-odds-ratio - SCORAD (calculator; dermatologia, alergia-e-imunologia, pediatria): Gravidade da dermatite atópica. https://github.com/Elucenia/tool-scorad - Sistema de Bethesda para citologia de tireoide (calculator; cirurgia-de-cabeca-e-pescoco, endocrinologia-e-metabologia, patologia): Risco de malignidade da PAAF de tireoide (2023). https://github.com/Elucenia/tool-sistema-de-bethesda-tireoide - Sódio corrigido na hiperglicemia (calculator; medicina-de-emergencia, medicina-intensiva, endocrinologia-e-metabologia): Katz (1,6) e Hillier (2,4). https://github.com/Elucenia/tool-sodio-corrigido-hiperglicemia - Escore SOFA (calculator; medicina-intensiva, medicina-de-emergencia, infectologia): Disfunção orgânica na sepse e na UTI. https://github.com/Elucenia/tool-sofa - sPESI (PESI simplificado) (score; pneumologia, cardiologia, medicina-de-emergencia): Prognóstico em 30 dias no TEP agudo. https://github.com/Elucenia/tool-spesi - Escala de Spetzler-Martin (score; neurocirurgia, radiologia-e-diagnostico-por-imagem, neurologia): Risco cirúrgico da malformação arteriovenosa. https://github.com/Elucenia/tool-spetzler-martin - SRQ-20 (Self-Reporting Questionnaire) (calculator; psiquiatria, medicina-de-familia-e-comunidade, medicina-do-trabalho): Rastreamento de transtornos mentais comuns. https://github.com/Elucenia/tool-srq-20 - STOP-Bang (score; anestesiologia, pneumologia, otorrinolaringologia): Rastreio de apneia obstrutiva do sono. https://github.com/Elucenia/tool-stop-bang - Superfície corporal e IMC (calculator; oncologia-clinica, cardiologia, pediatria): Mosteller e DuBois. https://github.com/Elucenia/tool-superficie-corporal - Superfície corporal queimada (Lund-Browder) (calculator; cirurgia-plastica, medicina-de-emergencia, medicina-intensiva): Percentual de SCQ corrigido pela idade. https://github.com/Elucenia/tool-superficie-corporal-queimada - Tamanho amostral para comparar duas proporções (calculator; medicina-preventiva-e-social): Amostra por grupo para ensaio clínico ou coorte. https://github.com/Elucenia/tool-tamanho-amostral-duas-proporcoes - Tamanho amostral para estimar uma proporção (calculator; medicina-preventiva-e-social, medicina-de-familia-e-comunidade): Amostra para prevalência com margem de erro definida. https://github.com/Elucenia/tool-tamanho-amostral-proporcao - Teste de Fagerström (score; medicina-de-familia-e-comunidade, pneumologia, psiquiatria): Grau de dependência à nicotina. https://github.com/Elucenia/tool-teste-de-fagerstrom - Teste diagnóstico (tabela 2×2) (calculator; medicina-preventiva-e-social, medicina-de-familia-e-comunidade, medicina-laboratorial): Sensibilidade, especificidade, VPP, VPN e razões de verossimilhança. https://github.com/Elucenia/tool-teste-diagnostico-2x2 - Taxa de filtração glomerular (CKD-EPI 2021) (calculator; nefrologia, clinica-medica, cardiologia): Estadiamento da doença renal crônica. https://github.com/Elucenia/tool-tfg-ckd-epi - TFG pediátrica (Schwartz à beira do leito) (calculator; pediatria, nefrologia): Filtração glomerular estimada em crianças. https://github.com/Elucenia/tool-tfg-de-schwartz - Escore TIMI (SCA sem supra de ST) (calculator; cardiologia, medicina-de-emergencia): Angina instável e IAM sem supra. https://github.com/Elucenia/tool-timi-sca - Gravidade da colecistite aguda (Tokyo 2018) (calculator; cirurgia-do-aparelho-digestivo, cirurgia-geral, gastroenterologia): Grau I, II ou III e conduta pelas diretrizes de Tóquio. https://github.com/Elucenia/tool-tokyo-2018-colecistite - UAS7 (Escore de Atividade da Urticária) (score; alergia-e-imunologia, dermatologia): Atividade da urticária crônica em 7 dias. https://github.com/Elucenia/tool-uas7 - VEF₁ e DLCO previstos pós-operatórios (calculator; cirurgia-toracica, pneumologia, cirurgia-oncologica): Risco da ressecção pulmonar pela contagem de segmentos. https://github.com/Elucenia/tool-vef1-dlco-pos-operatorio - VO₂ estimado, METs e capacidade funcional (calculator; medicina-esportiva, cardiologia): A partir do tempo no protocolo de Bruce. https://github.com/Elucenia/tool-vo2-e-mets - Volume prostático (elipsoide) (calculator; urologia, radiologia-e-diagnostico-por-imagem): Volume da próstata pelas três medidas. https://github.com/Elucenia/tool-volume-prostatico - Escore de Wells (embolia pulmonar) (calculator; pneumologia, medicina-de-emergencia, cardiologia): Probabilidade pré-teste de TEP. https://github.com/Elucenia/tool-wells-tep - Escore de Wells para TVP (score; angiologia, cirurgia-vascular, medicina-de-emergencia): Probabilidade clínica de trombose venosa profunda. https://github.com/Elucenia/tool-wells-tvp - Zonas de treino pela FC de reserva (Karvonen) (calculator; medicina-esportiva, cardiologia, medicina-fisica-e-reabilitacao): Frequência cardíaca-alvo pelo método de Karvonen. https://github.com/Elucenia/tool-zonas-de-treino-karvonen ## Faultline (open source tool by Felipe Guedes) A zero-dependency auditor for the boundaries in your codebase. One command reads your code for outbound calls without timeouts, retries without backoff, money endpoints without idempotency keys, swallowed errors, SQL and shell built from strings, secrets in code, cookies without flags and queries that forget the tenant, and tells you how to fix each one. Official page: https://fgxdev.com/faultline/ · PT-BR: https://fgxdev.com/pt/faultline/ · Source: https://github.com/thefgxdev/faultline · License: AGPL-3.0-or-later · Version: 0.1.1 Install and run: npx github:thefgxdev/faultline . --fail-on high Rules (22): Credential committed in source; Password or secret assigned as a literal; Outbound call without a timeout; Retry loop without backoff or jitter; Error caught and discarded; SQL built by string concatenation or interpolation; Shell command built from variables; eval or new Function on dynamic input; HTML injected from a variable; Math.random used for a token, code or password; Cookie set without httpOnly / secure / sameSite; JWT verified without pinning algorithms; CORS allows any origin with credentials; Money or irreversible endpoint without an idempotency key; Query in a multi-tenant codebase without a tenant filter nearby; SELECT * in application code; List query without a limit; target="_blank" without rel="noopener"; Dockerfile uses :latest or runs as root; .env file present and not git-ignored; No lockfile committed; HTTP server without a rate limiter in dependencies. **What is Faultline?** Faultline is a free, open source command-line tool that audits a codebase for boundary failures: outbound calls without timeouts, retries without backoff, money endpoints without idempotency keys, swallowed errors, SQL and shell commands built from strings, secrets committed in code, cookies without security flags, CORS misconfiguration and queries that forget the tenant. It prints every finding with the file, the line, the evidence and the fix. **Is Faultline free? Can I use it at work?** Yes. Running Faultline on any code, commercial or not, is free and creates no obligation. The AGPL-3.0 license only applies if you modify Faultline itself and distribute it or offer it as a service: then your modified source must be published under the same license. Closed modifications are available under a commercial license from the author. **Does Faultline send my code anywhere?** No. It runs locally, has zero dependencies, makes no network calls and has no telemetry. It reads your files and prints a report. **How is Faultline different from ESLint, Semgrep or Snyk?** ESLint checks style and language usage. Semgrep and Snyk are broad platforms with rule marketplaces, accounts and services. Faultline is one file-tree walk with 22 opinionated rules about the places where systems actually fail in production: timeouts, retries, idempotency, error handling, injection, secrets, auth and tenant isolation. It is meant to run before a human review and in CI, in seconds, with nothing to install. **How do I run Faultline in CI?** Add the GitHub Action (uses: thefgxdev/faultline@main with path and fail-on) or run npx github:thefgxdev/faultline . --fail-on high in any pipeline. The process exits with code 1 when a finding is at or above the chosen severity, and can write a Markdown report and JSON output. **What about false positives?** The rules are heuristics with context windows, tuned on real codebases and tested against fixtures of good and bad code. When a rule is wrong for your case, silence that line with // faultline-ignore or the path with .faultlineignore; the silence stays visible in code review. Report false positives with the smallest sample that reproduces them and they become test cases. **Which languages does Faultline support?** The rule set is JavaScript and TypeScript first (Node.js, Next.js, Express, Fastify, Prisma, Supabase and plain SQL strings), plus Dockerfiles, .env files, package.json and lockfiles. The file matchers already accept Python, Go, PHP and Ruby; rules for those boundaries are on the roadmap. **Who made Faultline?** Felipe Guedes, Software Engineer and Systems Architect based in Toledo, Paraná, Brazil, after ten years and 600+ systems built, reviewed or audited in 19 countries. The tool encodes the questions he asks in every software audit. ## NextFoot (football game, in development) Match simulation, player and team modelling, real-time state and a multiplayer-ready architecture, where latency and consistency are decided by design, not by luck. Built by Felipe Guedes as the hardest system he knows how to make feel simple. Page: https://fgxdev.com/nextfoot/ · PT-BR: https://fgxdev.com/pt/nextfoot/ · Status: in development, no public build and no release date. **What is NextFoot?** NextFoot is a football game built from the engine up by Felipe Guedes: deterministic match simulation, player and team modelling, real-time state and a multiplayer-ready architecture, rendered in WebGL in the browser. It is in development. **Can I play NextFoot?** Not yet. There is no public build and no release date. The project page and the author's profiles will announce a playable version when it exists. **What is the technology behind NextFoot?** TypeScript, a custom real-time engine with a deterministic simulation, multiplayer netcode with client prediction and server reconciliation, WebGL rendering in the browser, Node.js on the server and PostgreSQL for accounts, seasons and results. **Who is building NextFoot?** Felipe Guedes, Software Engineer and Systems Architect in Toledo, Paraná, Brazil, creator of Faultline and founder of ELUCENIA, after 600+ systems built, reviewed or audited in 19 countries. ## Articles (145) - [SLOs for a five-person team](https://fgxdev.com/articles/slo-for-a-five-person-team/) (2026-09-21; infra, reliability): You do not need an SRE org to have an SLO. You need one number, one window and one rule about what happens when you miss. - [Shipping is a habit](https://fgxdev.com/articles/shipping-is-a-habit/) (2026-09-19; career, engineering): Nobody decides not to ship. They decide to do one more thing first. How to build the habit that ends that. - [Designing a design system that survives its designers](https://fgxdev.com/articles/designing-a-design-system-that-survives/) (2026-09-18; fullstack, ux): A design system dies when the people who made it leave. The ones that live encode decisions in tokens and constraints, not in memory. - [I find the failure nobody finds. Here is the method.](https://fgxdev.com/articles/i-find-the-failure-nobody-finds/) (2026-09-16; audit, career): It is not intuition. It is a repeatable procedure that assumes the system is lying and goes looking for where. - [Design the failure before the feature](https://fgxdev.com/articles/design-the-failure-before-the-feature/) (2026-09-14; system-design, reliability): The happy path is the easy part. The system is everything that happens when it breaks, and that is what to design first. - [Testing non-deterministic systems](https://fgxdev.com/articles/testing-nondeterministic-systems/) (2026-09-11; ai, engineering): You cannot assert on an exact string from a model. You can assert on properties, distributions and invariants, and you should. - [The twelve questions I ask in every architecture review](https://fgxdev.com/articles/the-architecture-review-questions-i-always-ask/) (2026-09-09; architecture, audit): The same twelve questions, in the same order, on every system. They find most of what matters within a day. - [AI writes the code. You still own the system.](https://fgxdev.com/articles/ai-writes-the-code-you-still-own-the-system/) (2026-09-05; ai, engineering): The model can type faster than you. It cannot answer for what happens at 3 a.m. That part is still yours. - [The real cost of a retry](https://fgxdev.com/articles/the-cost-of-a-retry/) (2026-09-02; system-design, reliability): Retries add load exactly when the system is weakest. Backoff, jitter, budgets, idempotency and breakers keep them from amplifying. - [The slowest loop in science is not thinking](https://fgxdev.com/articles/the-slowest-loop-in-science-is-not-thinking/) (2026-08-31; science, ai): Everyone wants AI to make scientists think faster. The time is not lost in thinking. It is lost in waiting, and waiting is engineering. - [The whiteboard is where the system is actually built](https://fgxdev.com/articles/the-whiteboard-is-where-the-system-is-built/) (2026-08-30; system-design, architecture): The decisions that hurt later are made in the first hour, before any code. Here is what I draw and why. - [The AI feature nobody asked for](https://fgxdev.com/articles/the-ai-feature-nobody-asked-for/) (2026-08-28; ai, ux, business): The sparkle button was built for the board, not the user. Here is how to tell, and what to build instead. - [Where business rules should live, and where they end up](https://fgxdev.com/articles/where-business-rules-should-live/) (2026-08-26; architecture, engineering): The rule is written once in the spec and five times in the code. How it scatters, and how to bring it home. - [The CDN is the cheapest scaling decision you will make](https://fgxdev.com/articles/cdn-is-the-cheapest-scaling/) (2026-08-23; infra, performance): Before adding servers, ask which responses never needed to reach a server at all. - [Writing the audit report nobody wants to read, so that they do](https://fgxdev.com/articles/the-audit-report-nobody-wants-to-read/) (2026-08-21; audit, career): An audit is only as good as the fixes it causes. Structure, language and ranking decide whether the report gets acted on or filed. - [Why I keep choosing Next.js for products that have to last](https://fgxdev.com/articles/why-i-keep-choosing-nextjs/) (2026-08-19; fullstack, nextjs): Not because it is popular. Because it lets one team own the whole request path without inventing glue. - [A latency budget for LLM features](https://fgxdev.com/articles/llm-latency-budget/) (2026-08-15; ai, performance): The model is the slowest dependency you have ever added. Budget for it the way you budget for a database, per feature, per step. - [Boundaries are the product](https://fgxdev.com/articles/boundaries-are-the-product/) (2026-08-12; architecture, engineering): Features come and go. The boundaries you draw between them are what you will still be paying for in five years. - [You own the infrastructure whether you like it or not](https://fgxdev.com/articles/you-own-the-infra-whether-you-like-it-or-not/) (2026-08-09; infra, fullstack): Every line of application code runs on decisions you either made or accepted by default. Both are yours. - [What a latency histogram is trying to tell you](https://fgxdev.com/articles/reading-a-latency-histogram/) (2026-08-05; system-design, observability): Averages hide. Read the shape, then p50, p95 and p99, and learn why fan-out turns everyone else's p99 into your p50. - [The ORM is fine. Your queries are not.](https://fgxdev.com/articles/the-orm-is-fine-your-queries-are-not/) (2026-08-02; fullstack, data): Every slow database I audit has the same shape: an ORM blamed for queries nobody read. Read the SQL. Fix the pattern. - [The strangler fig, done honestly](https://fgxdev.com/articles/the-strangler-fig-done-honestly/) (2026-07-29; architecture): The pattern works. What kills it is pretending the old system will die on its own. It will not. - [Guardrails that do not ruin the product](https://fgxdev.com/articles/guardrails-that-do-not-ruin-the-product/) (2026-07-26; ai, ux): Most guardrails are a refusal bolted onto the end. Good ones are designed into the task so the user rarely meets them. - [Auditing a system you did not build, without offending the people who did](https://fgxdev.com/articles/auditing-a-system-you-did-not-build/) (2026-07-24; audit, career): An audit that produces resentment produces no fixes. Here is how I deliver hard findings and keep the team on my side. - [CI that tells you the truth](https://fgxdev.com/articles/ci-that-tells-you-the-truth/) (2026-07-22; infra, engineering): A green build that lies is worse than no build. Here is what it takes to make CI a signal you can act on. - [Idempotency is a business decision](https://fgxdev.com/articles/idempotency-is-a-business-decision/) (2026-07-21; system-design, engineering): Call it twice, get the same result. Sounds technical. Deciding what "the same" means is where the business signs. - [The hypothesis loop, and where software can shorten it](https://fgxdev.com/articles/the-hypothesis-loop/) (2026-07-18; science, ai): Science is a loop, not a line. Software should shorten the cheap turns and leave the expensive ones exactly as slow as truth requires. - [React state that belongs to the server](https://fgxdev.com/articles/react-state-that-belongs-to-the-server/) (2026-07-15; fullstack, nextjs): Most useState calls I audit are holding a copy of the database. Put that state where it lives and the bugs go with it. - [Hallucination is a system property, not a model bug](https://fgxdev.com/articles/hallucination-is-a-system-property/) (2026-07-11; ai, reliability): The model generates plausible text. Whether plausible becomes false in front of a user depends on the system you built around it. - [Why most microservices should have been modules](https://fgxdev.com/articles/why-most-microservices-should-be-modules/) (2026-07-08; system-design, architecture): The network boundary is a permanent tax. Four reasons a service earns it, and the signals that yours should have been a module. - [Say no to the rewrite](https://fgxdev.com/articles/say-no-to-the-rewrite/) (2026-07-06; career, architecture): The rewrite promises a clean slate and delivers a second system chasing the first. Five questions before you approve one. - [Containers are not isolation](https://fgxdev.com/articles/containers-are-not-isolation/) (2026-07-03; infra, security): A container is a process with a costume. Treat it as a security boundary and you will eventually regret it. - [Forms are distributed systems](https://fgxdev.com/articles/forms-are-distributed-systems/) (2026-07-01; fullstack, engineering): A form is two machines, an unreliable network and a user who clicks twice. Design it like the distributed system it is. - [Core Web Vitals are an architecture problem](https://fgxdev.com/articles/core-web-vitals-are-architecture/) (2026-06-30; fullstack, performance): You cannot fix LCP with a plugin. The metrics are measuring your data flow, your boundaries and your deploy. - [Why a software engineer cares about science](https://fgxdev.com/articles/why-a-software-engineer-cares-about-science/) (2026-06-28; science, career): Science is the oldest system for producing reliable knowledge. An engineer who ignores it is leaving the best design on the table. - [AI pair programming changed what I put on the whiteboard](https://fgxdev.com/articles/ai-pair-programming-changed-my-whiteboard/) (2026-06-26; ai, system-design): When the code is cheap, the whiteboard stops being about code. It becomes about failure, ownership and invariants. - [Domain events versus notifications: they are not the same thing](https://fgxdev.com/articles/domain-events-vs-notifications/) (2026-06-24; architecture): One is a fact the business cares about. The other is a poke. Confusing them is how event-driven systems rot. - [Faith, family and uptime](https://fgxdev.com/articles/faith-family-and-uptime/) (2026-06-21; life): What I hold above my work, and how those constraints made my systems more reliable, not less. - [Evaluating LLM features like an engineer, not a fan](https://fgxdev.com/articles/evaluating-llm-features-like-an-engineer/) (2026-06-19; ai, audit): If your quality signal is a demo and a feeling, you do not have a feature. You have a hope. Here is how I build evals. - [Rate limiting without hurting the users you want](https://fgxdev.com/articles/rate-limiting-without-hurting-good-users/) (2026-06-17; system-design): Token bucket, identity over IP, per-plan tiers, shadow mode: how to stop abuse without throwing 429s at paying customers. - [From papers to questions: what an instrument for research should do](https://fgxdev.com/articles/from-papers-to-questions/) (2026-06-14; science, ai): A summarizer gives you less to read. A research instrument gives you better questions to ask. The difference is the whole product. - [Postgres is enough, until the day it is not, and how to know that day](https://fgxdev.com/articles/postgres-is-enough-until-it-is-not/) (2026-06-12; infra, data): Most teams leave Postgres too early or too late. The signals that tell you which are measurable. - [Designing for deletion](https://fgxdev.com/articles/designing-for-deletion/) (2026-06-10; architecture, data): Every system can create a record. Few can delete one correctly. Design the delete path first and the rest gets simpler. - [Code review is risk review](https://fgxdev.com/articles/code-review-is-risk-review/) (2026-06-08; audit, engineering): Most reviews check whether the code is nice. The question that matters is what breaks, for whom, and how you get it back. - [TypeScript generics that help, and the ones that hurt](https://fgxdev.com/articles/typescript-generics-that-help-vs-hurt/) (2026-06-05; fullstack, typescript): A generic that relates an input to an output is a gift. A generic that appears once in a signature is a lie the compiler cannot catch. - [Designing APIs people can misuse safely](https://fgxdev.com/articles/designing-apis-people-can-misuse-safely/) (2026-06-02; system-design, engineering): Clients retry, duplicate, send stale data and call out of order. Design so the ordinary misuse is harmless. - [What AI cannot audit](https://fgxdev.com/articles/what-ai-cannot-audit/) (2026-05-30; ai, audit): Models are excellent at finding what looks wrong. The failures that matter are the ones that look right, and those still need a person. - [Multi-tenancy: the decision you cannot undo](https://fgxdev.com/articles/multi-tenancy-the-decision-you-cannot-undo/) (2026-05-27; system-design, architecture): Shared schema, schema per tenant or database per tenant: each is a different set of pains, and switching later means migrating everything. - [Evidence-first software](https://fgxdev.com/articles/evidence-first-software/) (2026-05-25; science, engineering): Every claim a system makes should be able to say why. Here is how I build software where the evidence comes before the conclusion. - [Zero-downtime migrations in practice](https://fgxdev.com/articles/zero-downtime-migrations-in-practice/) (2026-05-22; infra, data): The schema change is never the hard part. The lock, the backfill and the old code still running are. - [The full-stack engineer is a systems engineer](https://fgxdev.com/articles/the-full-stack-engineer-is-a-systems-engineer/) (2026-05-20; fullstack, career): Full stack was never about knowing two languages. It is about owning the failure wherever it happens. - [What a doctor taught me about systems](https://fgxdev.com/articles/what-a-doctor-taught-me-about-systems/) (2026-05-19; life, system-design): I met my father once, in his clinic. His way of diagnosing patients is how I audit systems today. - [The human in the loop is a design decision](https://fgxdev.com/articles/the-human-in-the-loop-is-a-design-decision/) (2026-05-15; ai, science): A human in the loop is not a checkbox for the compliance slide. It is an interface, a workflow and a budget you have to design. - [The shared library trap](https://fgxdev.com/articles/the-shared-library-trap/) (2026-05-13; architecture): The package that was created to avoid duplication is often the tightest coupling in the system. How to tell, and what to do instead. - [Details are not small](https://fgxdev.com/articles/details-are-not-small/) (2026-05-11; audit, life): A detail is a decision at small scale. The systems that fail, and the people we remember, are made of them. - [The first three alerts every system needs](https://fgxdev.com/articles/the-first-three-alerts/) (2026-05-08; infra, reliability): Before dashboards, before SLOs, before anything: three alerts that catch most outages and almost never lie. - [How I read a system I have never seen in one hour](https://fgxdev.com/articles/how-i-read-a-system-in-one-hour/) (2026-05-06; system-design, audit): The audit method I use to understand an unfamiliar system fast: entry points, data, money, and what happens when it is down. - [Error boundaries, end to end](https://fgxdev.com/articles/error-boundaries-end-to-end/) (2026-05-01; fullstack, reliability): An error boundary is not a React component. It is a decision about where failure stops, made at every layer from the click to the database. - [What a green build hides](https://fgxdev.com/articles/what-a-green-build-hides/) (2026-04-30; audit, infra): A passing pipeline proves the tests you wrote pass on the machine they ran on. Everything else is faith. Here is what to check. - [File uploads, done right the first time](https://fgxdev.com/articles/file-uploads-done-right/) (2026-04-28; fullstack, engineering): Never let a file pass through your application server. Sign, upload direct, verify async, and treat every filename as hostile. - [Observability for LLM applications](https://fgxdev.com/articles/observability-for-llm-apps/) (2026-04-24; ai, observability): Logs and latency are not enough. You need the prompt, the output, the cost and the quality per request, and the ability to replay it. - [Architecture decision records that actually get read](https://fgxdev.com/articles/architecture-decision-records-that-get-read/) (2026-04-22; architecture, engineering): Most ADRs are written once and never opened again. Here is the format and the discipline that make them survive. - [Self-taught is not alone-taught](https://fgxdev.com/articles/self-taught-is-not-alone-taught/) (2026-04-22; career): I never finished a degree, but I was never alone. How to find teachers when nobody assigns you one. - [Data provenance for scientists, explained by an engineer](https://fgxdev.com/articles/data-provenance-for-scientists/) (2026-04-18; science, data): A number you cannot trace is an opinion with decimals. What provenance means and how to build it without a PhD in databases. - [Streaming UI and what it costs the backend](https://fgxdev.com/articles/streaming-ui-and-what-it-costs-the-backend/) (2026-04-15; fullstack, nextjs): Streaming makes the page feel fast. It also holds connections open, fans out queries and hides slow paths. Budget for it. - [When not to use a language model](https://fgxdev.com/articles/when-not-to-use-a-language-model/) (2026-04-11; ai, system-design): A language model is a component with a cost, a latency and an error rate. Half the places I see one, a regex would have won. - [Designing for the slow dependency, not the dead one](https://fgxdev.com/articles/designing-for-the-slow-dependency/) (2026-04-09; system-design, reliability): Dead dependencies fail fast. Slow ones fill your pools and take you down. Bulkheads, deadlines, fallbacks and p99. - [Structured logs or nothing](https://fgxdev.com/articles/structured-logs-or-nothing/) (2026-04-05; infra, observability): A log line you cannot query is a log line you will read once, during an incident, too late. - [The modular monolith in practice](https://fgxdev.com/articles/modular-monolith-in-practice/) (2026-04-01; architecture): One deploy unit, hard internal boundaries. What it looks like in a real TypeScript codebase, and where teams cheat. - [The hidden cost of every dependency you add](https://fgxdev.com/articles/the-hidden-cost-of-a-dependency/) (2026-03-31; fullstack, audit): The install is free. The transitive tree, the upgrade cadence, the bundle weight and the supply chain risk are not. - [The ethics of building tools for medicine when you are not a doctor](https://fgxdev.com/articles/the-ethics-of-building-tools-for-medicine/) (2026-03-28; science, life): I am not a physician. I build tools that physicians and scientists will use. Here is the line I draw and why. - [When to add a message broker, and when you are avoiding a decision](https://fgxdev.com/articles/when-to-add-a-message-broker/) (2026-03-25; system-design): What a broker buys, what it costs, and the tell that you are using a queue to avoid deciding who owns the data. - [Prompt injection is input validation you forgot](https://fgxdev.com/articles/prompt-injection-is-input-validation/) (2026-03-20; ai, security): We solved SQL injection by separating code from data. Prompt injection is the same lesson in a system that cannot fully separate them. - [Types at the edges: validate once, trust everywhere](https://fgxdev.com/articles/types-at-the-edges/) (2026-03-18; fullstack, typescript): TypeScript types vanish at runtime. Put a real check at every edge, then let the compiler carry the proof inward. - [The first thing I check in any codebase](https://fgxdev.com/articles/the-first-thing-i-check-in-any-codebase/) (2026-03-15; audit, engineering): Before the architecture, before the tests, before the README: I find every place the data is written and count them. - [The cloud cost nobody budgets for](https://fgxdev.com/articles/the-cost-of-cloud-nobody-budgets/) (2026-03-13; infra, business): Compute is the line everyone watches. The bill grows in the lines nobody reads. - [Reversibility as an architecture goal](https://fgxdev.com/articles/reversibility-as-an-architecture-goal/) (2026-03-11; architecture, system-design): The best architecture decision is the one you can undo in a sprint. Optimize for that before optimizing for anything else. - [Building world-class systems from a small city in Paraná](https://fgxdev.com/articles/building-from-a-small-city/) (2026-03-08; career, life): Location matters. Here is what you have to do differently when you are never in the room. - [The prompt is not the spec](https://fgxdev.com/articles/the-prompt-is-not-the-spec/) (2026-03-06; ai, engineering): A prompt tells the model what to do today. A spec tells everyone what correct means. Teams keep confusing the two. - [Back-pressure is a feature, not a bug](https://fgxdev.com/articles/back-pressure-is-a-feature/) (2026-03-03; system-design, reliability): A system that says "not now" is healthier than one that says yes to everything and falls over. Build the refusal path. - [AI and the end of boilerplate, and what fills the space](https://fgxdev.com/articles/ai-and-the-end-of-boilerplate/) (2026-02-28; ai, career): Generating code is now cheap. The scarce skills moved to specification, verification and taste, and the career moved with them. - [DNS TTLs and other small things that bite at 3 a.m.](https://fgxdev.com/articles/dns-ttl-and-other-things-that-bite/) (2026-02-27; infra): The outages that hurt most are rarely architectural. They are a number someone set once and forgot. - [State machines for everything that matters](https://fgxdev.com/articles/state-machines-for-everything-that-matters/) (2026-02-24; system-design, engineering): Boolean flag soup hides contradictions. Explicit states, a transition table and a transitions log make bugs and reports obvious. - [What I look for in an engineer](https://fgxdev.com/articles/what-i-look-for-in-an-engineer/) (2026-02-22; career): Not the stack, not the years. Five habits that show up in the first hour and predict everything after. - [The deploy is part of the design](https://fgxdev.com/articles/the-deploy-is-part-of-the-design/) (2026-02-20; infra, system-design): If you cannot describe how a change reaches production safely, the design is not finished. - [The bug is always in the boundary](https://fgxdev.com/articles/the-bug-is-always-in-the-boundary/) (2026-02-19; audit, architecture): Inside a module, code is consistent with itself. Bugs live where two things with different assumptions meet. - [TypeScript as a design tool, not a linter](https://fgxdev.com/articles/typescript-as-a-design-tool/) (2026-02-17; fullstack, typescript): If your types only catch typos, you are using a tenth of the tool. Types are where I design the system first. - [How I review code a model wrote](https://fgxdev.com/articles/reviewing-ai-generated-code/) (2026-02-13; ai, audit): Same standards as any pull request, but the bugs live in different places. Here is where I look. - [Eventual consistency, explained to the product team](https://fgxdev.com/articles/eventual-consistency-explained-to-the-product-team/) (2026-02-11; system-design): Why the order shows up three seconds later, which flows must never lag, and how to design a UI that tells the truth. - [Cardiovascular data is a systems problem](https://fgxdev.com/articles/cardiovascular-data-is-a-systems-problem/) (2026-02-08; science, data): The heart produces more kinds of data than almost any organ. Making them agree is not a medical problem. It is an integration problem. - [The edge runtime: when it helps and when it just moves the problem](https://fgxdev.com/articles/edge-runtime-when-and-when-not/) (2026-02-05; fullstack, nextjs): Compute near the user is only fast if the data is near the compute. Otherwise you have added a round trip to every query. - [Reading legacy code like an archaeologist](https://fgxdev.com/articles/reading-legacy-code-like-an-archaeologist/) (2026-02-03; architecture, audit): Legacy code is not bad code. It is a record of decisions under pressure. Learn to read the layers before you dig. - [Agents need boundaries too](https://fgxdev.com/articles/agents-need-boundaries-too/) (2026-01-31; ai, architecture): An agent is a loop with permissions. Every rule I apply to services applies to it, with a larger blast radius. - [The API route that became a monolith](https://fgxdev.com/articles/the-api-route-that-became-a-monolith/) (2026-01-29; fullstack, nextjs, architecture): It started as one handler. Two years later it is a 900-line file that nobody dares to touch. Here is how to prevent that. - [Caching is a consistency decision in disguise](https://fgxdev.com/articles/caching-is-a-consistency-decision/) (2026-01-27; system-design, architecture): A cache is a copy, and every copy forces you to decide how stale a user is allowed to see the world. - [Validation is not a phase](https://fgxdev.com/articles/validation-is-not-a-phase/) (2026-01-24; science, engineering): Software teams learned that testing at the end does not work. Science tools need the same lesson: validate every step, or it is a draft. - [Load testing the right thing](https://fgxdev.com/articles/load-testing-the-right-thing/) (2026-01-22; infra, performance): Most load tests measure how fast the server returns a cached homepage. Production fails somewhere else. - [Contracts before code](https://fgxdev.com/articles/contracts-before-code/) (2026-01-20; architecture, engineering): Write the interface, the schema and the failure cases before the first implementation line. It is the cheapest design review you will get. - [A security review you can do in an afternoon](https://fgxdev.com/articles/security-review-in-an-afternoon/) (2026-01-17; audit, security): Not a penetration test. The seven checks that find most of what actually gets exploited, in four hours, with tools you already have. - [WebSockets, SSE or polling: a decision table](https://fgxdev.com/articles/websockets-vs-sse-vs-polling/) (2026-01-15; fullstack, system-design): Three ways to get live data to a browser. The right one depends on direction, scale and your hosting, not on which sounds modern. - [The outbox pattern in plain words](https://fgxdev.com/articles/the-outbox-pattern-in-plain-words/) (2026-01-13; system-design, data): Commit the event with the data, let a relay publish it, make consumers idempotent. The dual-write problem, solved in plain words. - [Research is a system with a throughput problem](https://fgxdev.com/articles/research-is-a-system-with-a-throughput-problem/) (2026-01-10; science, system-design): Scientists are not slow. The pipeline around them is. Here is how a systems engineer reads a research lab. - [Blue-green, canary and the honest rollback](https://fgxdev.com/articles/blue-green-canary-and-the-honest-rollback/) (2026-01-08; infra): Deployment strategies are only as good as the rollback they actually allow. Most teams overestimate theirs. - [Hexagonal architecture without the ceremony](https://fgxdev.com/articles/hexagonal-without-the-ceremony/) (2026-01-06; architecture): Ports and adapters is one good idea buried under a pile of folders. Keep the idea, drop the pile. - [The generalist advantage](https://fgxdev.com/articles/the-generalist-advantage/) (2026-01-05; career, fullstack): Failures live at the boundaries between layers, and the generalist is the one who can see the seam. - [Embeddings are not magic. They are a compression of your data.](https://fgxdev.com/articles/embeddings-are-not-magic/) (2026-01-03; ai, data): An embedding is a lossy summary chosen by someone else's training data. Treat it like a compression format, not an oracle. - [Server Components change where the boundary is](https://fgxdev.com/articles/server-components-change-where-the-boundary-is/) (2025-12-30; fullstack, nextjs, architecture): The old boundary was the API. The new one is a directive at the top of a file. Most teams have not noticed. - [The staging environment lie](https://fgxdev.com/articles/the-staging-environment-lie/) (2025-12-27; infra, engineering): Staging proves the code runs. It almost never proves the code runs in production, and pretending otherwise is expensive. - [What a new service really costs](https://fgxdev.com/articles/the-cost-of-a-new-service/) (2025-12-23; architecture, system-design): The code is the cheap part. This is the invoice that comes with every box you add to the diagram. - [The model will change. The interface should not.](https://fgxdev.com/articles/the-model-will-change-the-interface-should-not/) (2025-12-20; ai, architecture): Design the boundary around the model so that swapping providers, versions or prompts is a config change, not a rewrite. - [Why I still write SQL by hand](https://fgxdev.com/articles/why-i-still-write-sql-by-hand/) (2025-12-18; fullstack, data): The database is the most capable engine in your stack. A query builder hides that capability. I write the queries that matter myself. - [A graceful degradation checklist for product teams](https://fgxdev.com/articles/graceful-degradation-checklist/) (2025-12-16; system-design, reliability): Decide on purpose what users see when recommendations, search, payments or the database fail. A checklist for product teams. - [What clinicians taught me about requirements](https://fgxdev.com/articles/what-clinicians-taught-me-about-requirements/) (2025-12-14; science, career): I thought I knew how to gather requirements after hundreds of projects. Then I sat with people whose mistakes have a pulse. - [Secrets management for small teams that will grow](https://fgxdev.com/articles/secrets-management-for-small-teams/) (2025-12-11; infra, security): The .env file is fine on day one and a liability on day ninety. Here is the path that does not require a rewrite. - [The database is not your integration layer](https://fgxdev.com/articles/the-database-is-not-your-integration-layer/) (2025-12-09; architecture, data): Two systems sharing a database are one system with two deploy pipelines and no contract. Here is how to get out. - [Read the error handling first](https://fgxdev.com/articles/reading-the-error-handling-first/) (2025-12-08; audit, reliability): Catch blocks are where a team wrote down what it fears and what it chose to ignore. Read them before anything else. - [The cost of a token in production](https://fgxdev.com/articles/the-cost-of-a-token-in-production/) (2025-12-06; ai, business): The price per token is the smallest part of the bill. Context, retries and loops are where the money actually goes. - [Auth is not a library decision](https://fgxdev.com/articles/auth-is-not-a-library-decision/) (2025-12-04; fullstack, security): Pick the library last. First decide who the identity belongs to, where the session lives and what happens when it is revoked. - [The hot partition problem and why sharding does not save you](https://fgxdev.com/articles/the-hot-partition-problem/) (2025-12-02; system-design, data): Sharding spreads keys, not load. When one key is hot by itself, you need a different toolbox. - [The year I lost every contract](https://fgxdev.com/articles/the-year-i-lost-every-contract/) (2025-12-01; career, life): In 2020 every contract vanished in weeks. What I rebuilt had a different shape, and the shape is the useful part. - [Fine-tuning, retrieval or prompting: a decision I make weekly](https://fgxdev.com/articles/fine-tuning-vs-retrieval-vs-prompting/) (2025-11-29; ai): Three tools, three kinds of problems. Most teams pick by excitement. I pick by what changes and how often. - [A post-mortem template that produces change](https://fgxdev.com/articles/an-incident-postmortem-template/) (2025-11-27; infra, reliability): Most post-mortems are well-written and change nothing. The template decides which kind you write. - [The monolith is not the problem, the coupling is](https://fgxdev.com/articles/the-monolith-is-not-the-problem/) (2025-11-25; architecture): Splitting a tangled monolith into services gives you a tangled distributed system. Fix the coupling first. - [Six hundred projects later: what repeats](https://fgxdev.com/articles/six-hundred-projects-later/) (2025-11-22; audit, career): After 600+ projects built, reviewed or audited, the failures stopped surprising me. Here is the short list that keeps coming back. - [Feature flags as architecture](https://fgxdev.com/articles/feature-flags-as-architecture/) (2025-11-20; fullstack, engineering): A flag is a branch in your system that lives in production. Treat it like one, or it becomes the bug nobody can reproduce. - [The queue you did not know you had](https://fgxdev.com/articles/the-queue-you-did-not-know-you-had/) (2025-11-18; system-design): Your diagram shows one queue. Your system has twenty. Little's law explains why they all fill at the same time. - [Reproducibility is an engineering problem](https://fgxdev.com/articles/reproducibility-is-an-engineering-problem/) (2025-11-15; science, engineering): 'It worked on my laptop' is not a scientific result. The tools to fix that already exist, and they are boring. - [Observability is a product feature](https://fgxdev.com/articles/observability-is-a-product-feature/) (2025-11-13; infra, observability): The ability to answer "what happened to this user" is worth more than most items on your roadmap. - [API versioning is a relationship, not a number](https://fgxdev.com/articles/api-versioning-is-a-relationship/) (2025-11-11; architecture, engineering): The version in the URL is the least important part. What matters is who depends on you and what you owe them. - [RAG is a data pipeline with a language model at the end](https://fgxdev.com/articles/rag-is-a-data-pipeline/) (2025-11-08; ai, system-design): Most RAG failures I audit are ingestion and retrieval bugs wearing an AI costume. Fix the pipeline, then worry about the prompt. - [Caching in Next.js without surprises](https://fgxdev.com/articles/caching-in-nextjs-without-surprises/) (2025-11-06; fullstack, nextjs): Four cache layers, one rule: never cache what you cannot name, and never name what you cannot invalidate. - [Exactly-once is a promise nobody keeps](https://fgxdev.com/articles/exactly-once-is-a-promise-nobody-keeps/) (2025-11-04; system-design, engineering): Delivery and processing are different promises. At-least-once plus an idempotent consumer is the guarantee you actually get. - [Strict mode is the cheapest code review you will ever get](https://fgxdev.com/articles/strict-mode-is-the-cheapest-review/) (2025-10-30; fullstack, typescript): A reviewer who reads every line, on every keystroke, never tires and has no ego. It costs one line in tsconfig. - [Capacity planning on a napkin](https://fgxdev.com/articles/capacity-planning-on-a-napkin/) (2025-10-28; system-design): Back-of-the-envelope math that tells you whether you are building for 10 rps or 10,000, and which number you are wrong about. - [Small models, big leverage](https://fgxdev.com/articles/small-models-big-leverage/) (2025-10-25; ai, system-design): The frontier model is the wrong default. Most work in a real pipeline is narrow, repetitive and cheap to get right with a small model. - [A TypeScript monorepo that stays fast after year two](https://fgxdev.com/articles/a-typescript-monorepo-that-stays-fast/) (2025-10-23; fullstack, typescript): Monorepos do not get slow because they are big. They get slow because nobody drew the dependency graph on purpose. - [Conway's law is a tool, not a curse](https://fgxdev.com/articles/conways-law-is-a-tool/) (2025-10-21; architecture, career): Your system will mirror your org chart whether you like it or not. So draw the org chart on purpose. - [The pull request that looks fine](https://fgxdev.com/articles/the-pr-that-looks-fine/) (2025-10-19; audit, engineering): Small diff, green tests, clear description, quick approval. Here is the anatomy of the change that took production down anyway. - [Backups you have never restored are hypotheses](https://fgxdev.com/articles/backups-you-have-never-restored/) (2025-10-17; infra, reliability): A backup is a claim about the future. Only a restore turns it into a fact. - [Schema evolution without downtime](https://fgxdev.com/articles/schema-evolution-without-downtime/) (2025-10-15; system-design, data): Expand and contract, batched backfills, lock-aware migrations and feature-flagged cutovers: how to change a schema while it runs. - [Structured outputs are contracts](https://fgxdev.com/articles/structured-outputs-are-contracts/) (2025-10-11; ai, engineering): A schema turns a model from a text generator into a component you can integrate. Treat it with the seriousness of an API. - [Timeouts are the cheapest resilience you will ever buy](https://fgxdev.com/articles/timeouts-are-the-cheapest-resilience/) (2025-10-09; system-design, reliability): No timeout means infinite. Here is how I set connect, read and total budgets so one slow dependency cannot take the whole system down. - [Open science needs boring infrastructure](https://fgxdev.com/articles/open-science-needs-boring-infrastructure/) (2025-10-05; science, infra): Open science does not fail on ideals. It fails on storage, identifiers, uptime and the unpaid person keeping the server alive. - [Technical debt is a loan, and every loan has a name on it](https://fgxdev.com/articles/technical-debt-is-a-loan-with-a-name/) (2025-10-02; architecture, career): Debt is not the shortcut. It is the shortcut without an owner, a term and a payment date. Fix the metaphor and the backlog moves. - [From the print shop to the whiteboard](https://fgxdev.com/articles/from-the-print-shop-to-the-whiteboard/) (2025-10-01; career, life): Vinyl cutting, electronics repair and construction taught me the three habits I still use to design systems. ## Links - github: https://github.com/thefgxdev - linkedin: https://www.linkedin.com/in/fgxdev - x: https://x.com/thefgxdev - instagram: https://instagram.com/eufeguedes - memorial: https://pedromorettiguedes.com.br/ - email: contato@fgxdev.com - Portuguese (Brazil) version of the site: https://fgxdev.com/pt/ - Portuguese (Brazil) version of this file: https://fgxdev.com/llms-pt-br.txt - Father's memorial (Dr. Pedro Moretti Guedes, physician, 1935–2010): https://pedromorettiguedes.com.br/ # Full articles (English) ## SLOs for a five-person team 2026-09-21 · infra, reliability Service level objectives have a reputation for being a big-company thing. Error budgets, burn rates, policy documents, a dedicated reliability team to run the meetings. A five-person team reads about it and concludes, reasonably, that it is not for them. That conclusion is wrong, and the reason is that an SLO is not the machinery around it. It is a single sentence that ends an argument. The argument is the one every small team has: should we ship the feature or fix the flakiness? Without an SLO, that argument is decided by whoever is louder that week. With one, it is decided by a number. ## The sentence 'Ninety-nine point five percent of checkout requests succeed within two seconds, measured over thirty days.' That is an SLO. It names a user-facing thing, a definition of good, a target, and a window. Everything else in the SLO literature is elaboration of that sentence. Choosing the thing: pick the one user action whose failure hurts most. Not the API in general, not uptime, the checkout, the booking, the search. One thing. You can add a second later. Choosing the definition of good: a request is good if it returns a success status within a latency bound. Both conditions, because a slow success is a failure from the user's chair. The latency bound is whatever makes the product feel broken, usually one to three seconds for an interactive action. Choosing the target: look at the last thirty days of data and see what you already achieve. If it is 99.7 percent, set the target at 99.5. An SLO you already meet with a little margin is useful. An aspirational one is a wish. An SLO of 99.99 for a team of five is a wish with a pager attached. ## The budget The gap between the target and one hundred is the error budget. At 99.5 percent over thirty days, you may have roughly three and a half hours of full failure, or a much longer stretch of partial failure, before the objective is missed. That number is the whole point. It converts reliability from a feeling into a resource that is spent. When the budget is healthy, ship features, take risks, deploy on Friday if you want. When the budget is mostly spent, the next sprint is reliability work, and this is not a negotiation, because everyone agreed to the sentence in advance. That rule, agreed before the argument, is what makes an SLO work in a team with no reliability organization. ## What it takes to measure Less than people expect. You need a count of requests to the chosen action, a count of those that were good, and the ability to compute the ratio over a rolling window. If you already have structured logs or request metrics at the edge, this is one query. Put the ratio on a dashboard with the target drawn as a line. Add a single alert: if the budget is burning fast enough to be gone within a few days, tell someone. That is the entire implementation. Do not build burn-rate alerts with multiple windows on day one. Do not write a policy document. Do not create a dashboard with fifteen SLOs. One sentence, one graph, one alert, one rule. ## Where small teams go wrong - Choosing uptime of the server instead of success of a user action, so the SLO is green while users cannot check out. - Setting the target from ambition instead of from measured history, so the budget is always spent and the rule is always ignored. - Measuring at the application instead of at the edge, so timeouts at the load balancer never count as failures. - Skipping the rule about what happens when the budget is gone, so the SLO becomes a number nobody acts on. The value of an SLO for a five-person team is not the measurement. It is that the measurement makes a decision for you that you would otherwise have to fight about every two weeks. One sentence, agreed in a calm moment, replaces a dozen arguments in tense ones. That is worth an afternoon of setup at any team size. --- ## Shipping is a habit 2026-09-19 · career, engineering There are engineers who ship and engineers who almost ship. The gap between them is not talent. In my experience, across a lot of projects and a lot of teams, it is not even effort. It is a habit, and habits can be built deliberately. I want to describe the habit as precisely as I can, because "just ship it" is useless advice. Nobody decides not to ship. They decide to do one more thing first, and the one more thing has no end. ## Define done before you start The engineer who almost ships starts with an idea of what the feature should be and keeps refining it while building. Done is a feeling. The engineer who ships writes down, before the first line, what "done" means: which user can do which thing, verified how. Done is a sentence on a page. This sounds bureaucratic and takes about five minutes. It changes everything, because when the temptation comes to add one more thing, you can look at the sentence and see that the thing is not in it. It goes on the list for next time. The current work ships. ## Ship the smallest thing that is true I learned this on a repair bench long before I wrote code. You do not fix everything on a board and then test. You fix one fault, power it up, and see if that was the one. Small change, immediate feedback, next change. In software the same rhythm looks like this: the smallest slice that a real user can touch goes to production first. Not the whole feature. The first true piece of it. A form that saves but does not yet validate. An endpoint that returns real data for one client. Each slice teaches you something the plan did not know, and each slice is a deployment, so the deploy path stays warm and unscary. ## Make the release boring Teams that ship rarely have a release ritual, a release night, a release person. Teams that ship often have none of those things, because releasing is so routine it has no ceremony. If your release is an event, you will avoid it, and avoidance is how the almost-shipped feature grows for three months. The practical work here is infrastructure, and it is worth doing before the feature: one command to deploy, one command to roll back, a health check that tells you within a minute whether the deploy is healthy. With that in place, shipping a small change is less stressful than not shipping it. ## Stop polishing what nobody has seen Perfectionism is the main enemy of this habit, and I say that as someone who is described as detail-obsessed. The difference is where the detail goes. I am obsessive about the correctness of what ships: the migration is reversible, the timeout exists, the error is handled. I try not to be obsessive about the completeness of what ships. Correct and small beats complete and unreleased. The test I use: has a real user seen this? If not, the polish is speculative. You are making decisions about what matters without the one source of information that would tell you. ## The compounding part Here is why it is a habit and not a technique. Each time you ship, the next ship is a little easier. The deploy path is warmer, the definition of done is more natural, the fear is smaller. Each time you delay, the opposite happens. The branch grows, the merge gets scarier, the feature accumulates scope, and the next delay is more likely. After enough repetitions the habit becomes identity. You are the person who ships. People bring you work because they know it will come out the other side. That reputation, in my career, has been worth more than any technology I know. So start this week. Pick the thing you have been almost finishing. Write one sentence that defines done. Cut everything not in the sentence. Ship it. Then do it again on Thursday. --- ## Designing a design system that survives its designers 2026-09-18 · fullstack, ux I have audited products whose design system was beautiful, documented and dead. The designer who built it left, the engineer who maintained it moved teams, and within a year every new screen was built with one-off styles because nobody knew the rules anymore. The system was in their heads. When they left, it left with them. A design system that survives is one where the decisions live in the code, in a form the compiler and the linter can enforce, and where a new engineer can build a correct screen on their first week without asking anyone. ## Tokens are the contract, not the components The lasting part of a design system is not the button. It is the set of named values the button is made of: spacing, color roles, type scale, radii, shadows, motion durations. Those are the tokens, and they are the layer that outlives any component library or framework. The discipline is semantic naming. A token called blue-500 tells you nothing about when to use it. A token called surface-interactive or text-danger tells you exactly when, and it lets you change the actual color for dark mode or a rebrand without touching a single component. Every raw value in a component file is a decision someone will have to rediscover. I export tokens as one source of truth that generates the CSS variables, the TypeScript constants and the design tool's library. When they drift, and they will if maintained separately, the system is already dying. ## Constrain the props The second survival mechanism is that components accept decisions, not styles. A button takes a variant and a size from a closed union. It does not take a className that lets anyone override its padding, and it does not take a style prop. That feels restrictive on day one. It is what keeps the product visually coherent on day seven hundred, when forty engineers have shipped screens. TypeScript enforces this for free. A variant typed as a union of four strings cannot become a fifth without a pull request to the system, and that pull request is where the design conversation happens. Escape hatches exist, but they are named, ugly and grep-able, so that an audit can count them. Accessibility lives in the primitives or nowhere. Focus rings, keyboard handling, ARIA roles and contrast are properties of the base components, tested once, inherited everywhere. A system that leaves accessibility to each screen has decided not to have it. ## Documentation that cannot rot Written documentation drifts. Examples that compile do not. Every component ships with usage examples that are real code, rendered in a catalog and run in CI, so that a breaking change to a component fails the build of its own documentation. The catalog is also where design and engineering meet. When a designer proposes a new pattern, the question is whether it can be built from existing tokens and components. If yes, it ships. If no, the system grows on purpose, with a name and a rule, rather than by accident in a feature branch. ## Ownership and the deprecation path A system with no owner has no future. Someone, by name, decides what enters and what leaves. That is a part-time job in a small company and a team in a large one, but it cannot be nobody. Nothing is removed without a deprecation period, a lint warning that points to the replacement, and a codemod when the change is mechanical. A system that breaks its consumers trains them to fork it, and a forked design system is two dead design systems. The test I run on any design system is simple. Give a new engineer a screen from the product's roadmap and one day. If they can build it without a raw color value, a custom padding or a question to a designer, the system is alive. If they cannot, it is already gone, whether or not the people who built it are still there. --- ## I find the failure nobody finds. Here is the method. 2026-09-16 · audit, career Clients hire me to audit systems that other people already reviewed. The pull requests were approved, the tests were green, a consultant signed off. And I find something. Often something serious. People call it a gift. It is not. It is a procedure, and I can write it down. ## Assume the system is lying Every system tells a story about itself: the README, the architecture diagram, the comments, the dashboard. The first rule is to treat all of it as a hypothesis, not a fact. The diagram shows a queue between two services. Does the queue exist? Is it used? Does anything consume from it? I have found queues drawn on diagrams that were never deployed, and retries described in comments that were never implemented. The story is where people believe the system is. The failure lives in the gap between the story and the code. ## Trace one real request, end to end, by hand Not a happy path in a test. A real request from production logs, with a real ID, followed through every service, every database write, every external call, every retry, until it either finishes or is lost. I do this with a notebook open and I write down every place a decision is made: a branch, a timeout, a catch block, a default value. It takes hours. Every serious failure I have ever found showed up on this list as a decision nobody remembered making. ## Ask what happens the second time The first execution of anything usually works. The second one is where systems break. What happens if this webhook arrives twice? If this job restarts halfway? If the user double-clicks? If the deploy runs the migration again? I ask this question at every write, every external call and every scheduled task. Idempotency is where the largest class of silent data corruption hides, and almost nobody tests it because the tests run once. ## Read the data, not just the code Code describes what the system is supposed to do. The database records what it actually did. I query production, read-only, and look for the impossible: orders with a negative total, users with two active subscriptions, events with a timestamp before their parent's, statuses that the state machine says cannot coexist. Every impossible row is a bug that already happened. The code review missed it because the code looked right. The data does not care how the code looked. While I am there, I read every catch block, because each one is a place where the system decided to keep going without the information it just lost. A catch that logs and continues after a failed payment confirmation is not error handling. It is a decision to lose money quietly. ## Reproduce before you report A finding without a reproduction is an opinion. Before anything goes in the report, I trigger it: a script that sends the duplicate webhook, a query that shows the impossible rows, a request with a timestamp in a different timezone. If I cannot make it happen, I say so and rank it lower. This is the step that turns 'I think there is a problem' into 'here is the problem, here is how to see it, here is what it costs'. It is also the step that makes people fix things instead of arguing about them. The people who built the system are not less capable than me. They are closer to the story. They wrote the diagram, they remember the intent, and their brain fills in the gap between intent and code without noticing. I have no intent to protect. I have a list of places where systems fail, built across 600+ projects, and I check every one of them, every time, even when the system looks fine. Especially when it looks fine. --- ## Design the failure before the feature 2026-09-14 · system-design, reliability When I review a design document, I skip the happy path. I look for the section that describes what happens when the payment provider times out, when the queue fills, when two users click the same button at the same moment. Most documents do not have that section. That is the whole problem. ## The feature is the easy part A feature is a sequence of steps that works when everything works. Any competent engineer can write it. The system is what happens around that sequence: the retries, the partial writes, the duplicate events, the user who refreshes mid-request. That is where the real design lives, and it is where I find the failures nobody else finds, because nobody looked. Here is the habit I recommend: before writing the first line of the feature, write the failure list. Not a risk register with probabilities. A plain list of concrete sentences that start with "what happens when". ## The questions I ask before the code exists - What happens when this call times out after the write already committed? - What happens when the same request arrives twice? - What happens when the downstream service is up but ten times slower than usual? - What happens when this job runs at the same time as itself? - What happens when the user closes the tab between step two and step three? - Who notices, and how long does it take them? Six questions. On a checkout service they take twenty minutes on a whiteboard. Answering them after the feature is in production takes weeks, because every answer now touches data that already exists. ## A concrete example Take a clinic scheduling app. The feature: a patient books a slot. The happy path is a form, a validation, an insert. Now the failures. Two patients book the same slot within the same 200 milliseconds. The confirmation email provider is down. The patient double-clicks. The insert succeeds but the response is lost, so the client retries. Each of these forces a design decision. The double booking needs a unique constraint on (doctor, slot), not a check in application code. The email needs to be queued, not sent inline, so a broken provider does not block bookings. The double click and the lost response need an idempotency key so the second attempt returns the first booking instead of creating another one. None of this is exotic. But notice that all four decisions change the schema, the API contract, or the infrastructure. If you discover them after launch, you are migrating a live table with patient data in it. ## Failure design is cheaper than failure handling There is an asymmetry I keep repeating to teams. Designing a failure mode costs minutes. Handling it once it happens in production costs the incident, the data repair, the customer conversation, and the fix, which is now constrained by everything you already shipped. I am not asking anyone to build for every failure. I am asking them to decide, explicitly, which ones they accept. "If the email provider is down, bookings still work and the email goes out later" is a design. "If two people book the same slot, the second one gets an error" is a design. "We did not think about it" is not. ## What this looks like in a design review When I review, I want to see three things next to every feature. First, the list of failure sentences. Second, for each one, either the mechanism that handles it or an explicit line saying it is accepted and why. Third, the way someone will find out when it happens: a log line, a metric, an alert. If a document has those three things, the feature part is almost always fine. If it does not, the feature part is usually fine too, and the system is not. The feature is what the customer asked for. The failures are what the customer will remember. --- ## Testing non-deterministic systems 2026-09-11 · ai, engineering The first test I wrote for a model-backed feature asserted that the output equaled a string. It passed once. The second run produced a synonym and the test failed, and the third run produced the original again. I had written a flaky test on purpose without realizing it. Testing systems with a model inside requires giving up exact equality and picking up a different toolkit, one that engineers who work with distributed systems and randomized algorithms already know. The model is not untestable. It is testable the way a load balancer is testable: on properties, over many runs, with tolerances. ## Separate what is deterministic Most of the code around a model is ordinary code. Prompt assembly, retrieval, parsing, validation, routing, fallbacks. All of that gets ordinary unit tests with the model replaced by a stub that returns fixed outputs, including malformed ones. This is where most bugs live, and it is where tests are cheap and exact. I put the model behind an interface specifically so this layer can be tested without it. A team that tests "the AI feature" end to end and nothing else has a slow, flaky test suite and no idea which layer broke. ## Property tests for the model itself For the model call, I assert on properties of the output, not its content. The output parses as the schema. Required fields are present. The classification is one of the allowed labels. The summary is shorter than the input. The extracted date appears in the source text. The answer cites a passage that exists in the retrieved set. Each property is a real bug class I have seen in production, and each is checkable without knowing the exact wording. These run against the live model on a small fixed set, and they fail loudly when the provider changes something. ## Evals as statistical tests Quality is a distribution, so I test it like one. An eval set of a few hundred real inputs with expected outputs or a scoring rule, run against the current implementation, produces a score. The test is that the score does not drop below a threshold, and that it does not drop significantly compared to the last run. A single input flipping from right to wrong is noise. Twenty flipping is a regression. I run these on every prompt change and every model change, in CI, with the results stored so the trend is visible. This is the only test that tells you whether the feature got worse. ## Seeds, temperature and replay Where the provider allows it, I set temperature to zero and a fixed seed for tests, which removes most of the variance without removing all of it. I do not rely on it. What I rely on is recorded traces: real production requests with their real outputs, replayed against a new version and diffed. The diff is reviewed by a person or scored by a model. It is not a pass or fail test in the traditional sense, it is a change report, and it has caught more surprises than any assertion. ## What a green build means When a build passes, it means the deterministic code is correct, the model output satisfies the structural properties, and the quality score on the eval set has not regressed. It does not mean the feature is right for every input. Nothing means that. What the test suite gives you is the ability to change the prompt, the model or the retrieval on a Tuesday afternoon and know within ten minutes whether you broke something measurable. For a non-deterministic system, that is the whole game. --- ## The twelve questions I ask in every architecture review 2026-09-09 · architecture, audit After enough architecture reviews, you stop improvising. The systems differ; the questions do not. These are the twelve I ask on every review, in the order I ask them, and roughly what a bad answer looks like. They are not clever. They work because they are asked every time, on every system, with no exceptions for systems that look fine. ## The questions - What is the one operation this system exists to do, and can you show me its path end to end? - Where does money, or the thing that stands in for money, change hands, and what makes that step atomic? - What happens when the database is unavailable for thirty seconds? - What happens when the same request arrives twice? - Which piece of data is duplicated across two stores, and what keeps the copies aligned? - Which component would you be most afraid to deploy on a Friday, and why? - How does a request get authenticated, and where is that decision enforced other than at the edge? - Show me the last three incidents. What did each one teach, and what changed? - If I delete a customer, what is left behind? - What is the oldest dependency in production, and who owns upgrading it? - Which decision here would be most expensive to reverse, and was it made on purpose? - Who can explain the system without the person who built it in the room? ## Why these twelve The first question calibrates everything else. If the team cannot trace the main path without opening five repositories and a wiki, the system is already more complex than its owners can hold in their heads, and every other answer will be partial. Questions two through five are about correctness under stress. They are the ones that find the failures nobody else finds, because they ask about the boundaries: the transaction that spans services, the retry that creates a duplicate order, the cache that is stale by design but nobody wrote down for how long. A confident answer is fine. 'I think it is handled' is where I start reading code. Question six is a shortcut. Engineers always know which component is fragile. They rarely get asked. The answer usually names the place the next incident will come from, and it is free. Seven is about defense in depth. A system that checks permissions only at the gateway has one lock on the front door and every internal room open. Any internal call, script or misconfigured route bypasses it. Eight is about learning. Three incidents with three durable changes is a healthy system. Three incidents with 'we restarted it' is a system that will have the same incident a fourth time. Nine through eleven are about the future. Deletion reveals what owns what. The oldest dependency reveals how upkeep actually works. The expensive-to-reverse decision reveals whether the architecture was chosen or accumulated. Twelve is about the bus factor, and it is the one that most often changes my recommendation. A sound architecture that lives in one person's head is a risk, not an asset. ## How I use the answers I do not score them. I write down each answer, including the pauses, and I mark which ones were confident, which were uncertain and which were wrong. Wrong answers, the ones where the team said 'that is handled' and the code says otherwise, are the most valuable output of the review, because they show where the team's model of the system and the system itself have diverged. Then I prioritize by blast radius, not by elegance. A stale cache in a reporting dashboard is a note. A non-idempotent payment endpoint is the first item, in bold, with a date. ## What I do not ask I do not ask what framework they use, whether it is microservices or a monolith, or what the diagram looks like. Those are answers to questions nobody asked. I have seen beautiful diagrams for systems that could not survive a thirty-second database outage, and ugly monoliths that handled every one of these questions well. The list is yours to use. Ask all twelve, in order, on your own system, and write down where you hesitated. That hesitation is the review. --- ## AI writes the code. You still own the system. 2026-09-05 · ai, engineering A model can now write most of the code I used to write by hand. I use it every day and I am not going back. But something in the way people talk about it worries me: they describe the code as the product. It never was. The product is the system, and the system is a set of decisions about boundaries, failure, data, cost and accountability. The model does not own any of those. You do. ## The code was never the hard part In the projects I have audited, the failures that hurt were rarely in a function that was written badly. They were in a function that was written correctly for the wrong assumption. A retry that was safe on a read and catastrophic on a payment. A cache that was fine until two tenants shared a key. A queue consumer that assumed an ordering nobody guaranteed. None of those are typing problems. A model that writes flawless TypeScript will happily reproduce every one of them, because the assumption lives outside the code, in the head of whoever specified the task. ## What ownership actually means Owning a system means you can answer four questions without opening the editor. What happens when this dependency is slow? What happens when this runs twice? Who is allowed to see this data, and where is that enforced? What does this cost per unit of work at ten times the load? If you can answer those, the model can write the code and you can review it against the answers. If you cannot, the model will make a choice for you, silently, and it will be the average choice from its training data, not the right choice for your product. ## Where the model helps and where it does not The model is superb at the parts that are already decided. Given a schema, it writes the repository. Given a contract, it writes the client. Given a state machine, it writes the transitions and the tests. It is weak, in a way that is dangerous because it looks confident, at the parts that are not decided yet: which state machine, which contract, which invariants matter, what the failure should look like to the user. Those are the parts I still put on a whiteboard before I open a prompt. ## My working rules - I write the invariants before I ask for code. If I cannot state what must never happen, I am not ready to generate anything. - Every generated module gets the review a junior engineer's code would get, with extra attention to error handling and boundaries, because that is where models guess. - Generated tests do not count as coverage until I have read them and tried to break them. - The model never picks the data model, the consistency model or the auth model. It implements them. - If I could not have written it myself, slower, I do not ship it. I might use it to learn, but not to ship. ## The audit test Here is the test I apply to any team that tells me AI writes most of their code. I pick a module and ask the engineer who merged it to explain what it does under partial failure. Not what the code says, what the system does. When they can, the AI is a multiplier and the codebase is usually in good shape. When they cannot, I have found the failure nobody else found, and it is never in the code. It is in the ownership. The model gave them speed and they spent it on not understanding. That trade is available to everyone now, and it is the worst one on the menu. Use the model. Use it a lot. Then stand in front of the whiteboard and be able to draw what it built, because when it breaks at 3 a.m., the model is not on call. You are. --- ## The real cost of a retry 2026-09-02 · system-design, reliability A retry looks free. The call failed, so you call again. What could it cost? Quite a lot, as it turns out. Retries are the most common way I have seen a small incident turn into a full outage, and most teams add them without ever deciding what they should cost. The reason is simple. A retry adds load exactly when the system is least able to handle it. A dependency times out because it is overloaded. Your client responds by sending the same request again. Now the dependency is more overloaded. Multiply that by every caller, and you have a retry storm. ## The multiplication nobody calculates Picture three layers: a browser, an API gateway, a service that talks to the database. Each layer retries three times on failure. When the database gets slow, one user action can become 3 x 3 x 3 = 27 database calls. Twenty-seven. The database was already struggling with one. That is the arithmetic that turns a slow query into a dead cluster. It is not exotic. It is the default behaviour of most HTTP clients stacked on top of each other, each one configured by a different person who never saw the whole picture. I look for this in almost every audit, and I almost always find it. Nobody set out to build a 27x amplifier. They just added "retries: 3" in three separate config files. ## Backoff, jitter and budgets The first fix is exponential backoff: wait 100 ms, then 200, then 400, then 800. The second is jitter, meaning you randomise those waits. Without jitter, every client that failed at the same instant retries at the same instant, and you have simply scheduled the next spike. The third fix is the one most teams skip: a retry budget. Instead of allowing three retries per request, you allow retries to be at most some fraction of total traffic, say 10 percent. When the failure rate climbs past that, retries stop. The budget guarantees that a bad minute costs you 1.1x load, not 27x. Backoff and jitter shape the retry. The budget caps it. You want all three. ## Retry the right things, on the right errors A retry is only safe if the operation is idempotent. Retrying a read is fine. Retrying "charge the card" when the first attempt timed out after the charge succeeded is a double charge. If an operation is not idempotent, either make it idempotent with a request key or do not retry it automatically. Let a human or a reconciliation job decide. The error code matters too. Retry on timeouts, on 503, on connection resets: transient conditions that may resolve in seconds. Do not retry a 400 or a 404 or a 422. The request was wrong. Sending it again with the same payload is wasted load and, worse, it hides a bug behind a wall of noise. ## The brake: circuit breakers Retries are the accelerator. A circuit breaker is the brake. When failures exceed a threshold, the breaker opens and calls fail immediately, with no retries, for a cooling period. Then it lets a trickle through to test whether the dependency has recovered. Without a breaker, a retry policy will keep pushing load into a dependency that is on its knees. With a breaker, the dependency gets the one thing it needs to recover: silence. ## The retries you cannot see The most dangerous retries are the ones the server does not know about. Mobile apps that retry on their own. Browser code with a fetch wrapper somebody wrote two years ago. A service mesh with retries turned on by default. An SDK with built-in retries underneath your explicit retries. When you tune retries, you are tuning one layer. The system retries at every layer. Before you decide a policy, trace a single failing request from the user's finger to the disk and count the attempts. The number is almost always bigger than anyone in the room expected. A retry is a bet that the second attempt will succeed and that the cost of trying is small. Make that bet explicitly, with backoff, jitter, a budget, an idempotency guarantee and a breaker. Otherwise you are not retrying. You are amplifying. --- ## The slowest loop in science is not thinking 2026-08-31 · science, ai There is a fantasy version of AI for science where the model has the idea and the human nods. I do not believe in it, and more importantly, I do not think it addresses the actual problem. When I trace where time goes in a research project, thinking is the fast part. The slow part is everything a thought has to wait for before it can be tested. ## Where the time actually goes Waiting for access to a dataset that already exists somewhere. Waiting for a collaborator to send the version of the file they meant. Waiting for a cleaning script to be rewritten because the previous student took it with them. Waiting for the compute job that failed at hour six because of a path. Waiting for the literature search to get through the two hundred abstracts that looked relevant and were not. Waiting for the figure to be regenerated after the reviewer asked for one different threshold. Waiting for the review itself. In a software system I would draw this as a trace and see one long bar of idle time with a few short bars of actual work. That is not a thinking problem. It is a throughput problem, and it responds to different tools. ## What AI is actually good at here Language models are excellent at the tasks that sit between thoughts: reading a large corpus and extracting what is relevant to a precise question, restating a vague idea as a testable one, drafting the analysis code for a well-specified test, explaining a statistical method the researcher has not used before, translating between the vocabulary of two subfields that describe the same thing differently. Every one of those is a wait that becomes short. None of those is the thinking. The researcher still decides what question matters, still judges whether the extracted evidence is credible, still owns the result. The model removed friction, not judgment. ## What AI is bad at, and should stay bad at Deciding that a hypothesis is true. Generating a number that looks like a result without the analysis that produces it. Filling a gap in the evidence with a plausible sentence. I have spent enough time evaluating language model features in production to know that fluency and correctness are independent properties. A tool for science has to be built so that fluent output is never mistaken for validated output. That is a design decision, made upstream of the model, and it is where most of the work is. ## The engineering underneath The reason waiting dominates is not that scientists are disorganized. It is that the infrastructure underneath them was never designed as a system. Data has no stable interfaces. Code has no reproducible environment. Results have no provenance. Every handoff loses state. AI on top of that is a faster way to produce confusion. AI on top of a designed system, with versioned data, reproducible pipelines and evidence attached to every claim, is a genuine acceleration of the parts that deserve to be fast. So my order of operations is always the same: fix the plumbing, then add the intelligence. It is less impressive in a demo and it is the only way I know to produce something a scientist can trust. ## Keep the thinking slow I want to be explicit about this because the market pushes the other way. The slowness of real scientific thought is not a defect. Sitting with a result that does not fit, arguing with a colleague, sleeping on it, coming back to the data with a different question: that is where discoveries come from, and it should not be optimized. What should be optimized is the six weeks between having the question and being able to look at the answer. Those six weeks are engineering. They are mine to shorten. --- ## The whiteboard is where the system is actually built 2026-08-30 · system-design, architecture I have watched many systems get built, and the pattern is consistent: the decisions that matter were made in the first hour, usually standing in front of a whiteboard or a shared diagram, long before anyone opened an editor. Everything after that was implementation of those decisions, good or bad. ## Code is the cheapest thing to change People think code is expensive and diagrams are cheap, so they rush through the diagram to get to the code. It is backwards. Code is the most changeable artifact in a project. A refactor costs a day. A schema decision costs a migration. A boundary between two services costs a contract, a deployment coupling, and a year of coordination. The data model and the boundaries are what the whiteboard decides, and they are the things you will not be able to change later without pain. ## What I actually draw I do not draw architecture diagrams with pretty boxes. I draw four things, in this order. First, the data. What are the nouns, who owns each one, and which one is the source of truth. If two boxes both think they own the customer record, I have already found the first bug. Second, the writes. Every arrow that changes state. Reads are forgiving; writes are where consistency, ordering, and idempotency live. For each write arrow I ask what happens if it succeeds and the response is lost. Third, the boundaries. Where does one team, one deploy, one database end and another begin. Every boundary is a place where a partial failure can happen, so I want as few as the problem actually needs. Fourth, the failure paths, drawn in a different colour: the timeout, the retry, the queue that backs up. If the diagram has no red lines, it is a sales diagram, not a design. ## The questions a whiteboard answers that code cannot Code answers "does it work". The whiteboard answers "is this the right shape". Concretely: is there a single table that every feature writes to, which will become the hot spot? Is there a synchronous chain of four calls where one slow link stalls the user? Are there two sources of truth for the same fact? Is there a report that needs data from three services that will never be consistent at the same moment? These are shape problems. You cannot see them in a pull request, because a pull request shows one path. You see them in a diagram, because a diagram shows all paths at once. ## How to run a good whiteboard session Keep it small: two to four people, one of whom knows the business rules cold. Start with a concrete user story, not with components. Draw the data first, then walk the story through it, arrow by arrow. Every time someone says "and then it just calls", stop and ask what "just" is hiding. Write the decisions down as sentences, not as a picture, because the picture will be lost and the sentences will be read. One more rule: when someone proposes a component, ask what problem it solves that the current drawing does not. Half of the queues, caches and services proposed in these sessions are solutions looking for a problem, and a whiteboard is the cheapest place to say no. ## The part people skip The session ends with a list of things you decided not to handle. That list is as valuable as the diagram. Six months later, when something breaks, it tells you whether the break was a surprise or a known trade-off. Systems built without that list treat every incident as a mystery. Systems built with it treat most incidents as scheduled debt. I have audited enough systems to say this plainly: I can usually tell, from the schema and the boundaries alone, whether the whiteboard hour happened. When it did, the code is boring and the incidents are small. When it did not, the code is clever and the incidents are not. --- ## The AI feature nobody asked for 2026-08-28 · ai, ux, business There is a button in a lot of products now, usually with a sparkle icon, that opens a chat box. Nobody asked for it. It was built because a competitor had one, because a slide needed the letters AI on it, because a budget was allocated with those letters in the title. I have been asked to review the architecture of features like this many times, and the architecture is rarely the problem. The problem is that there was no user with a job to do at the start, so there is no way to know whether the feature did it. ## How to tell The signs are consistent. The feature is described by its technology, "an AI assistant", not by the outcome, "cuts the time to file a claim in half". The success metric is usage of the feature itself, not improvement in a metric the business already cared about. The entry point is a floating button, unconnected to any existing workflow. And when you ask what happens when the model is wrong, the answer is a shrug, because nobody has traced the output into a consequence. A feature with real stakes has an error budget. A feature built for the board has a demo script. ## What it costs The cost is not only the model bill, though that is real and it scales with usage of a feature that produces no value. The cost is the trust. Users try the sparkle button, get a fluent answer that is slightly wrong about their own data, and conclude that the product's AI is unreliable. That conclusion sticks, and it poisons the next feature, the one that might have been good. I have seen teams ship the useful feature a year later and fight an uphill battle against the reputation of the useless one. The first AI feature a product ships sets the expectation for all of them. ## Start from a job The alternative is boring and it works. Find a task the user already does, that is repetitive, that involves reading or writing or classifying, and that has a measurable outcome. Triage, drafting, extraction, search over their own documents, filling a form from a source. Build the model into that workflow, at the point where the user already is, with the output in the format they already use. Measure the outcome the user cares about: time, error rate, completion. If the number moves, you have a feature. If it does not, you have learned something for the cost of a small model and a few weeks, which is cheaper than a sparkle button that nobody uses and everybody remembers. ## Invisible is fine Some of the best model features I have built are not visible as AI at all. A field that is pre-filled correctly. A queue that is sorted by what actually needs attention. A search that finds the thing on the first try. Users do not thank the model. They just do the task faster and stop complaining about the old way. The instinct to make the AI visible comes from the same place as the sparkle button: it is for the people watching, not the people using. If the user does not need to know a model was involved, do not tell them, and spend the interface budget on making the result easy to verify and correct instead. ## The question to ask Before building any model feature, I ask one question and I insist on a specific answer: which existing number gets better, by how much, and how will we know by when. If the room cannot answer, the feature is for the room. Build it anyway if the politics demand it, but build it small, and put the real budget into the one with an answer. The board will forget the sparkle button. The users will not forget the thing that saved them an hour a day. --- ## Where business rules should live, and where they end up 2026-08-26 · architecture, engineering Ask a team where the rule 'an order above a certain amount needs manager approval' lives and watch what happens. The backend engineer points at a service. The frontend engineer points at a form validator. The data person points at a SQL view. The ops person points at a cron job that sends the approval email. They are all right, and that is the problem. Business rules are the most valuable code in a system and the most likely to be scattered. Everything else, the plumbing, the framework glue, the infrastructure, is replaceable. The rules are the reason the software exists. And in most codebases I audit they live in at least four places that disagree with each other. ## Where they end up The scatter follows a pattern. A rule is born in the domain layer, or wherever the team keeps its core logic. Then the frontend needs to show an error before the user submits, so the rule is copied into a validator. Then a report needs to count approvals, so the rule is encoded in SQL. Then a notification needs to fire, so a scheduled job re-derives who needs approval. Then an import script needs to bypass it for legacy data, so it gets a flag. Five copies. Each written by a different person under a different deadline. When the threshold changes, three copies are updated, one is missed, and the fifth is in a script nobody remembers. The worst place I find rules is inside the database, as triggers and stored procedures, invisible to anyone reading the application code. The second worst is the frontend, because it is the only copy the user actually experiences, and it silently defines the product regardless of what the backend thinks. ## Where they should live A rule should have exactly one authoritative home, and that home should be plain code with no framework attached: a function or a small class that takes domain values and returns a decision. No HTTP, no ORM, no date library that hides the current time. Testable with a unit test that runs in milliseconds. Everything else consults that home. The API handler calls it before persisting. The frontend either calls an endpoint that runs it or, if latency matters, runs the same code shipped as a package built from the same source. The report does not encode the rule; it reads the decision the rule already recorded, a column like requires_approval written at the moment the order was placed. That last point matters more than it looks. Recording decisions, not just data, is how you stop rules from being re-derived. If the order row says it needed approval, the report, the notification job and the audit log all read that. Nobody recomputes it, so nobody gets it wrong when the threshold changes and old orders should keep their old outcome. ## The checks that reveal scatter When I audit for this, I use three checks. - Grep for the constant. Pick a business number, a threshold, a rate, a limit, and search the whole repository including SQL and frontend. Every hit outside one file is a copy. - Change the rule on paper and trace the diff. If the change touches more than one module and one test file, the rule has escaped. - Ask for a decision's history. If nobody can say why a specific order was approved last March without re-running today's logic, decisions are being re-derived instead of recorded. ## How rules escape, and how to stop it Rules escape for good reasons. The frontend wants instant feedback. The report wants to run without calling an API. The batch job runs at night when the service might be down. Each copy solves a real problem, so telling people 'do not duplicate' does not work. What works is giving each of those needs a sanctioned path. For instant feedback: ship the rule as a pure package the frontend can import. For reports: record decisions at write time. For batch jobs: have them call the service, and make the service reliable enough for that, which is a separate problem you should solve anyway. And name the home. In every codebase I set up there is a directory that is visibly the domain, with a short README: rules live here and nowhere else. It does not stop scatter by itself. It gives the reviewer who spots a threshold in a SQL view a place to point at. Systems where rules have one home are the ones where a product change is a one-line pull request. Systems where rules are scattered treat every change as an archaeology project. The difference is not the framework or the language. It is whether someone decided, early, where the truth lives. --- ## The CDN is the cheapest scaling decision you will make 2026-08-23 · infra, performance When a product starts to feel slow under load, the reflex is to add compute. Bigger instances, more replicas, a faster database tier. Those work, and they cost money every hour forever. The cheaper move, and the one I check first in every performance review, is to ask how many of the requests hitting the origin should never have arrived there. For most web products, the answer is most of them. Static assets, public pages, product listings, documentation, images, API responses that change once an hour. All of it can be served from an edge node near the user, and once it is, the origin sees a fraction of the traffic and the user sees a fraction of the latency. ## What a CDN actually buys Three things, and they are separate. First, distance: an edge node a few milliseconds from the user instead of an origin on another continent. For a page with a dozen assets, that is the difference between a second and a fifth of a second before anything renders. Second, offload: every cache hit is a request the origin does not serve, which means fewer instances, fewer database connections, and a smaller blast radius when the origin has a bad day. Third, absorption: a burst of traffic, whether a launch or an attack, hits the edge first, and the edge is built to take it. The bill for all of this is usually a fraction of the compute it replaces, and on many platforms the first terabytes are free. It is hard to find a better return in infrastructure. ## The part people get wrong The CDN is trivial for static files. The value, and the risk, is in caching dynamic responses, and that comes down to headers. A response with no cache headers is treated conservatively, which means the CDN does nothing. A response with a long max-age and a private user's data in it is a data leak served at the speed of light. So the discipline is per route. Public and slow-changing: cache at the edge with a max-age of minutes to hours, and add a stale-while-revalidate window so users never wait for a refresh. Public and fast-changing: short max-age, seconds to a minute, which still absorbs a burst. Personalized: no edge caching, ever, marked private, and the CDN should still terminate TLS and compress. The mistake I find most often is a cookie or authorization header that varies the response without being declared, so one user's cached page is served to another. Vary headers exist for this, and they are not optional. ## Invalidation, honestly Cache invalidation is famous for being hard. It is less hard if the cache keys are chosen carefully. Assets get a content hash in the filename and a max-age of a year, so they are never invalidated, only replaced. Pages get a short max-age and a tag, so a content change can purge everything with that tag in one call. API responses get an ETag so the CDN can revalidate cheaply instead of refetching. What does not work is 'purge everything on deploy'. It works technically and it throws away the entire benefit at the moment of highest risk, when a new version is starting up under full traffic. ## The checklist I run - Every asset has a content hash in its path and a one-year max-age. - Every public route declares an explicit cache policy; no route relies on defaults. - Every personalized route is marked private and never varies silently on a cookie. - Cache hit ratio is measured, and anything under eighty percent on public traffic is investigated. - Purge is by tag or path, tested in staging, and never 'everything'. Run that checklist on a product that has never had a CDN in front of it, and the origin traffic typically drops by more than half. That is scaling you did not pay compute for, and it is available before the first bigger instance is ever ordered. --- ## Writing the audit report nobody wants to read, so that they do 2026-08-21 · audit, career I have written hundreds of audit reports and I know exactly what happens to a bad one. It gets opened, scrolled, and forwarded to someone else with 'thoughts?'. Nobody reads page nine. The critical finding on page nine stays in production. So I stopped writing reports that prove how much I found and started writing reports that get things fixed. The difference is structure, and it is learnable. ## The first page is the whole report Whoever opens the document has five minutes, maybe less. The first page has to work alone: what I looked at, what I found that matters, in what order, and what to do first. Three findings, at most five, each in two sentences: what breaks and what it costs. If the reader closes the document after page one and fixes those three things, the audit succeeded. Everything after page one is supporting evidence for someone who wants it. ## Rank by consequence, not by category Security people rank by a vulnerability score. Engineers rank by how hard the fix is. Neither is what the person deciding needs. I rank by what happens if it is not fixed, in the terms the business already uses: loses customer data, loses money, loses availability, loses trust, and then everything else. A missing index that makes the checkout time out under load ranks above a theoretically exploitable header, because one is happening every Friday and the other requires an attacker who has not shown up. The ranking is the most important editorial decision in the document, and I spend as much time on it as on any finding. ## Every finding has the same four parts What: one sentence, concrete, naming the component. Evidence: how I saw it, reproducibly, with the query, the request or the screenshot. Impact: what it does to a user, a customer or the company, in plain words. Fix: the next step, with an effort estimate honest enough to plan around. No finding without evidence, because evidence ends arguments. No finding without a fix, because a problem without a next step is just anxiety. If I cannot propose a fix, I say so and rank it lower. ## Write in the reader's language The report is not for me and it is not for other auditors. It is for an engineering lead who has to decide what to do this sprint, and for a founder or director who has to decide what to fund. So I write 'a duplicate webhook creates a second charge on the customer's card' and not 'the endpoint lacks idempotency guarantees'. The technical detail goes in the evidence section, where the engineer who will fix it can find it. The sentence at the top has to be understood by someone who has never seen the codebase and never will. An audit that only lists problems reads as an attack, and people defend against attacks instead of acting on them. Every system I have ever looked at does something well, and I say so, specifically, near the top. Not to soften the blow. Because it is true, and because a team that hears its strengths named accurately trusts the same author when the weaknesses are named. It also tells them what not to break while fixing everything else. ## The test of a good report Six weeks after delivery I ask one question: which findings were fixed? If the answer is the top three, the report worked. If the answer is 'we are still discussing it', the report failed, no matter how thorough it was. I have found failures nobody else found, on systems that many capable people had reviewed. That is worth nothing if the finding dies in a document nobody read past the first scroll. Writing the report so that it gets read is not a soft skill on top of the audit. It is the last, and most important, step of the audit itself. --- ## Why I keep choosing Next.js for products that have to last 2026-08-19 · fullstack, nextjs People assume I pick Next.js because it is popular. That is not it. I pick it for products that have to survive five years, three teams and at least one pivot, and the reason is structural, not fashionable. I have reviewed hundreds of codebases built on frameworks that were the right call on day one and a liability by year two. The pattern that kills them is not the framework. It is the glue: the custom build pipeline, the hand-rolled data fetching layer, the bespoke routing conventions that only the original author understood. ## One request path, one owner What I actually want from a web framework is that a single engineer can trace a request from the browser to the database and back without changing repositories, languages or mental models. Next.js with the App Router gives me that. A page is a Server Component. It can read from the database directly, with no API layer in between, and hand serialized props to the client parts that need interactivity. That collapses a whole category of bugs I used to find in audits: the frontend and the backend disagreeing about the shape of the data. When both live in the same TypeScript program, the compiler catches the disagreement before a user does. Server Actions do the same thing for mutations. A form submits to a function that runs on the server, typed end to end, with no route handler to keep in sync. I still treat every action as a public POST endpoint, because that is what it is, and I validate its input like one. But the plumbing is gone. ## Defaults that age well The defaults matter more than the features. Since Next.js 15, fetch calls are not cached unless you say so. Pages are dynamic unless you opt into caching with an explicit directive. That is the correct default for a product that changes, because stale data is a bug you cannot reproduce and a surprise cache is the worst kind of surprise. Streaming with Suspense is another default I rely on. The shell of a page renders immediately and slow parts arrive when they are ready. That is not a performance trick, it is a design constraint that forces you to decide what is critical on every screen. File-based routing sounds trivial until you inherit a project with a 2,000-line route table. Conventions you cannot argue about are conventions you do not have to document. ## What I do not love I am not a fan. Fans do not audit. The caching model has changed more than once and every change created a class of projects that are half migrated. The build output can be opaque when something goes wrong, and debugging a hydration mismatch at 2 a.m. teaches you humility. The framework also makes it easy to blur the line between server and client code. That line is where most of the serious bugs I find live: secrets imported into client bundles, database calls inside components that someone later marks with `'use client'`. I enforce that line with the `server-only` package and with lint rules, not with trust. And the vendor question is real. Next.js runs fine on a plain Node.js server behind a reverse proxy, and I have deployed it that way many times, but some of the more advanced features are smoother on the hosting platform built by the same company. You should know that before you commit, and decide with your eyes open. ## The decision, honestly So here is the trade-off I accept. I get one language, one type system, one deployable unit and conventions strong enough to survive turnover. I pay with a framework that moves faster than I would like and that rewards discipline at the server-client boundary. For a weekend prototype, use whatever you enjoy. For a product that will process money, medical records or anything a company depends on, I want the boring path: TypeScript everywhere, Server Components by default, a relational database, and a framework whose conventions will still be legible to the engineer who replaces me. That is why I keep choosing it. Not because it is exciting. Because it lets the excitement go into the product instead of the plumbing. --- ## A latency budget for LLM features 2026-08-15 · ai, performance Before language models, the slowest thing in most of my systems was a database query on a cold cache, measured in tens of milliseconds. A model call is measured in seconds, and the tail is measured in tens of seconds. That is a different regime. You cannot bolt a component like that onto a request path designed for milliseconds and expect the user experience to survive. You have to budget for it, and the budget has to start from the user, not from the model. ## Where the time goes A model call has three parts. Time to first token, which is dominated by the size of the input and the queue at the provider. Generation time, which is roughly linear in output tokens. And everything around the call: retrieval, the prompt assembly, validation, retries. Agentic features multiply all of this by the number of steps. When a team tells me a feature is slow, I ask for a trace with those segments separated, because the fix for a slow input is different from the fix for a long output, and the fix for a chatty loop is different from both. ## Set the budget from the user, not the model The budget is the maximum time a user will tolerate for this action, in this context, before the feature feels broken. For an inline suggestion that is under a second. For a search-like answer it is a few seconds with visible progress. For a background report it can be minutes, as long as the user is not waiting on a spinner. I write that number down first and then design backwards: what can fit, what must be streamed, what must move off the request path entirely. A feature designed forwards from what the model takes ends up with a spinner and a churned user. ## Levers Once the budget exists, the levers are concrete. Send fewer input tokens: shorter prompts, tighter retrieval, summarized history. Ask for fewer output tokens: structured outputs instead of prose, a short answer with an expandable explanation. Use a smaller model for the steps that do not need a frontier one. Parallelize independent calls instead of chaining them. Cache anything repeated, including the model's own answers to identical requests. And set timeouts that reflect the budget, with a fallback that is useful, because a request that takes twenty seconds and then fails has cost the user more than an immediate honest error. ## Streaming and perceived latency Streaming does not make the model faster. It makes the wait feel shorter, because the user sees progress instead of nothing, and that is worth a lot. But streaming has a cost on the backend: the connection stays open, the server holds state, and every layer in between must pass chunks through without buffering. It also changes what you can validate, because you cannot check a schema on half a JSON object. My rule is to stream to the person and not to the program: text a human reads is streamed, structured output a system consumes is delivered whole, after validation. ## Measure p95, not the demo The demo runs at p50 on a quiet afternoon. Production runs at p95 during the provider's busiest hour, with a retry. I measure the full budget per feature at the tail, with the model call broken out from the rest, and I alert on the tail, not the average. A feature whose median is two seconds and whose p95 is eighteen is not a two-second feature. It is an eighteen-second feature that is sometimes fast, and the users who hit the eighteen seconds are the ones who write the reviews. The budget is a promise, and the tail is where promises are kept or broken. --- ## Boundaries are the product 2026-08-12 · architecture, engineering When I audit a system, I do not start with the features. I start with the lines. Where does one thing end and another begin? Who is allowed to call whom? What can change without waking up three other teams? The answers tell me more about the next two years of that company than any roadmap. Most teams think they are shipping features. They are shipping boundaries. The feature is what the user sees this quarter. The boundary is what every future feature will have to negotiate with. ## A feature is temporary, a boundary is not A checkout button gets redesigned four times. The decision that checkout, inventory and payments share one database table with a status column survives every redesign. It survives because nobody sees it, nobody owns it, and moving it requires touching everything at once. This is why I say the boundary is the product. It is the part of the system with the longest half-life. Code inside a well-drawn boundary can be rewritten in a week. Code that crosses a badly drawn boundary cannot be rewritten at all without a migration project that nobody wants to fund. The test I use is simple. Pick any module and ask: if I deleted this tomorrow and rebuilt it from scratch, what else would break? If the answer is 'only its callers, through a contract', the boundary is real. If the answer is 'we would have to check', the boundary is decorative. ## What a real boundary looks like A real boundary has four properties, and I check every one of them in a review. - It has a single, explicit entry point: a function, a module export, an HTTP endpoint, a queue topic. Not seven. - It owns its data. Nobody else reads or writes those tables, files or keys directly. - It can fail without taking its neighbors down. A timeout on one side is an error on the other, not a hang. - Its contract is written down somewhere that is not the implementation: types, a schema, an OpenAPI file, a test suite that a stranger could read. Notice that none of these require microservices. A folder in a Next.js project with one index.ts export and a private data access layer satisfies all four. A service with three public endpoints that also lets other services query its Postgres directly satisfies none of them. ## Where boundaries get drawn wrong The most common mistake is drawing boundaries around technology instead of around change. 'The API layer', 'the database layer', 'the UI layer'. Those are boundaries between kinds of code, not between kinds of decisions. A change to how discounts work will touch all three layers, so the layers protect nothing. The second mistake is drawing them around the org chart of the day. A boundary that exists because two teams did not want to talk to each other in 2023 will still be there in 2028 when those teams have merged, split and merged again. The third is drawing them too early. A boundary is a bet that two things will change independently. If you do not know yet how the domain moves, you are guessing. I would rather keep two things together for six months and split them with evidence than split them on day one and spend a year building bridges back. ## The question I ask product managers When a product manager brings me a feature, I ask one question before estimating anything: which boundaries does this cross? If it lives inside one, it is cheap and safe. If it crosses two, it needs a contract change and a coordinated deploy. If it crosses four, the feature is really an architecture change wearing a feature's clothes, and it should be priced like one. This question changes conversations. Product people are not stupid. When they can see that 'add a discount code at checkout' crosses pricing, cart, payments and reporting, they start asking whether the boundaries are in the right place. That is exactly the conversation an architect should want. Draw the lines on purpose. Write down why each one is there. Revisit them when the reason disappears. Everything else in the codebase is replaceable. The lines are what you are actually building. --- ## You own the infrastructure whether you like it or not 2026-08-09 · infra, fullstack There is a sentence I hear in almost every audit: 'we do not really do infra, we just deploy to the platform'. I understand the feeling. Nobody wants to spend their week on DNS records and IAM policies. But the sentence is false, and the falseness costs money. If your product runs somewhere, you own the infrastructure. You may not have written a Terraform file. You may have clicked through a dashboard and accepted every default. Those defaults are now your architecture. When they fail, the customer does not call the platform. The customer calls you. ## Defaults are decisions A managed Postgres with the default connection limit is a decision. A serverless function with the default timeout is a decision. A container with no memory limit is a decision. You did not make them on purpose, but you will debug them on purpose, usually at the worst possible hour. The trap is that defaults are tuned for the demo, not for your traffic. The default connection pool is small enough to exhaust with a modest Next.js deployment where every serverless instance opens its own connections. The default log retention is short enough that the evidence of last week's incident is already gone. The default backup window may be fine, or may be exactly when your batch job runs. You do not need to change every default. You need to know which ones you are accepting and why. I keep a plain list per project: every managed service, every default I kept, and the number that matters (limit, timeout, retention). It takes one afternoon and it has saved me weeks. ## The full-stack engineer's real stack I call myself full stack, and I mean it literally. The stack does not end at the ORM. It includes the runtime, the network between the runtime and the database, the load balancer, the DNS, the certificate that expires in ninety days, the CDN that caches a response you thought was private. Every one of those layers can produce a bug that looks like an application bug. A request that times out at the load balancer after thirty seconds looks like a slow query. A stale DNS entry looks like a flaky service. A CDN caching an authenticated page looks like a session leak. If you cannot read those layers, you cannot find those bugs, and someone will tell you the code is fine while the product is broken. ## What owning it actually means Owning infrastructure does not mean running your own servers. I use managed services on nearly everything, and I recommend the same to most teams. Owning it means four concrete things. - You can draw the request path from the browser to the database and back, including every hop that is not your code. - You know the three or four limits that will be hit first under load, and roughly at what traffic. - You can restore the system from nothing in a documented amount of time, and you have done it at least once. - You know what it costs per month and which line will grow fastest. If any of those is missing, the platform owns you, not the other way around. ## The cheap version For a small team this is not a big investment. One document with the request path and the limits. One environment file that lists every external dependency. One rehearsed restore. One monthly look at the bill. A few hours a month, and the difference between a team that is surprised by its infrastructure and a team that is merely inconvenienced by it. The engineers I trust most are not the ones who know the most about Kubernetes. They are the ones who never say 'that is not my layer'. Everything the customer touches is your layer. Accepting that early is much cheaper than discovering it during an outage. --- ## What a latency histogram is trying to tell you 2026-08-05 · system-design, observability An average latency is the least useful number on a dashboard. It answers a question nobody asks: what would each request take if the slow ones shared their pain with the fast ones. Nobody experiences the average. Users experience one request at a time, and the slow ones are the ones they remember. A histogram answers a better question: how are requests actually distributed? Learning to read one is one of the highest return skills in operations, and it takes an afternoon. ## Read the shape before the numbers When I open a histogram I ignore the percentile table for a moment and look at the shape. A healthy service usually shows a single hump with a long right tail. That's normal: most requests are fast, some are slow for boring reasons. Two humps mean two code paths. Almost always it's cache hit versus cache miss: one cluster at 5 ms, another at 80 ms. Sometimes it's a warm instance versus a cold start on a serverless platform. Sometimes it's a query that uses an index for most inputs and does a sequential scan for a few. A bimodal histogram is an invitation to name both paths and decide whether the slow one is acceptable. A hump with a small bump far to the right, at a suspiciously regular interval, is often the garbage collector or a periodic job holding a lock. The shape tells you where to look before you have a single number. ## What p50, p95 and p99 each tell you p50, the median, is the typical experience. If p50 moves, your normal path changed: a deploy, a bigger payload, a slower dependency for everyone. p95 is the experience of your engaged users. Someone who makes twenty requests in a session will hit a p95 request once. If p95 doubles while p50 is flat, something is wrong for a minority: one shard, one region, one customer with a large account. p99 is where the system's weakest parts show up: lock contention, retries, timeouts to a dependency, a noisy neighbour. It's the number the on-call engineer should watch, because incidents start there and spread down. The gap between them is a signal on its own. A p50 of 20 ms and a p99 of 40 ms is a tight, predictable service. A p50 of 20 ms and a p99 of 2 s is a service with a hidden failure mode. ## Tail amplification is why p99 matters more than you think Here's the number that changes how people design. If a page fans out to 10 backend calls in parallel, and each has a 1% chance of hitting its own p99, the chance that at least one call is slow is 1 minus 0.99 to the tenth power, roughly 10%. The page latency is the slowest call. So one page in ten experiences a backend p99, even though each backend is behaving perfectly by its own numbers. With 100 fan-out calls, which is normal in a search or feed service, that becomes 63%. Most pages hit a p99 somewhere. This is why large systems obsess over tails: their p50 is built from everyone else's p99. The fixes are structural: hedged requests, tighter timeouts with a fallback, fewer calls per page, and caching the calls with the worst tails. ## How the buckets work, and what to alert on Prometheus-style histograms don't store every observation. They store counts in cumulative buckets with fixed boundaries: le=0.005, le=0.01, le=0.025 and so on. Percentiles are then interpolated between bucket edges. That has two consequences worth knowing. First, your p99 is only as precise as the buckets near it. If your boundaries jump from 1 s to 2.5 s, every p99 in that range will read as an interpolated value that may mean nothing. Set boundaries around the latencies you actually care about, especially near your SLO. Second, percentiles from buckets can be aggregated across instances, unlike precomputed percentiles, which cannot be averaged. That's why you want histograms, not summaries, when you have many replicas. For alerting, avoid alerting on p99 directly; it's too noisy at low traffic. Alert on the fraction of requests slower than your threshold over a window, which is a plain ratio of two bucket counters, and on error budget burn rate. A single slow request at 3 a.m. is not an incident. Five percent of requests over 500 ms for ten minutes is. The histogram was always trying to tell you this. The average was just talking over it. --- ## The ORM is fine. Your queries are not. 2026-08-02 · fullstack, data At some point in every project's life, someone proposes removing the ORM because the database is slow. I have been asked to review that proposal many times. I have agreed with it once. The ORM is almost never the problem. The problem is that the ORM made it easy to write queries nobody ever read, and the query patterns those unread queries produce are the same ones that would be slow in hand-written SQL. Fix the patterns and the ORM is fine. Remove the ORM without fixing the patterns and the team will write the same slow queries by hand, with more bugs. ## The patterns that are always there The first is N+1. A list of fifty orders, and for each one a query for the customer. The ORM hid the loop behind a property access, and the log shows fifty-one queries per page. The fix is to declare the relation in the query so it becomes one join or two batched selects. Every ORM has a way. Nobody used it because nobody looked at the log. The second is over-fetching. Select everything from a table with a JSON column and a text column, to render a list that shows the name. The wire carries megabytes per request and the database reads pages it did not need. Select the columns you render, and the ORM will happily generate that. The third is filtering in application code. Fetch a thousand rows, filter to twenty in JavaScript. The database has an index for that. It never got the chance to use it. The fourth is the missing index, which no ORM can add for you. A where clause on a column with no index is a sequential scan, and it is fast at ten thousand rows and a catastrophe at ten million. The ORM did not cause it. The absence of a query plan review did. ## Read the SQL My first step on any database complaint is to turn on query logging for one hour of production traffic and sort by total time. Not by average, by total: a two-millisecond query that runs a million times is the one hurting you, and it is invisible in a slow query log with a one-second threshold. Then I take the top five and run them with the plan explained. In nearly every case, the top five account for most of the load, and at least three of them are one of the four patterns above. That is a day of work to fix, and it usually ends the conversation about removing the ORM. I keep that logging on permanently in a sampled form. A query that was fine last quarter and is slow now changed because the data changed, and you want to know the week it happened, not the month a customer noticed. ## What the ORM is actually for An ORM earns its place by making the common query safe and typed. Insert a record and get a typed object back. Update by id. Load an aggregate with its children. Those are ninety percent of the queries in a product, and hand-writing them is where the SQL injection and the mistyped column live. The other ten percent are reporting queries, complex aggregations, window functions, anything with a common table expression. Those I write in SQL, by hand, with a typed result, and I keep them in a module next to the ORM code. The two coexist without conflict. What does not work is forcing the ORM to express a query it was not built for, because the result is unreadable and slow, and it makes the ORM look guilty. ## The one time I agreed The case where removing the ORM was right involved a data pipeline that processed millions of rows in batches, and the ORM's per-row object allocation was the bottleneck. That is a throughput problem, not a query problem, and it is rare. If your application serves web requests, it is not your case. Read the SQL your ORM generates. Add the indexes the plan asks for. Batch the relations. Then decide whether the tool was ever the issue. It was not. --- ## The strangler fig, done honestly 2026-07-29 · architecture The strangler fig is the migration pattern everyone recommends and almost nobody finishes. You put a facade in front of the legacy system, route new functionality to the new system, move old functionality across piece by piece, and one day the legacy system has nothing left and you switch it off. The pattern is sound. I recommend it constantly. But I have audited enough half-strangled systems to know where the honesty breaks down, and it is always the same three places. ## The facade is a system too The first dishonesty is treating the facade as a temporary shim. A proxy that routes some paths to the old system and some to the new one is a piece of infrastructure with its own failure modes, its own latency, its own deploy pipeline and its own need for observability. If it goes down, both systems are unreachable. If it routes wrong, the bug is invisible to both teams because each one sees only its own traffic. So the facade gets built properly or the migration fails at month three. It needs routing rules in configuration, not in code. It needs per-route metrics so you can see what percentage of traffic has moved. It needs the ability to send the same request to both systems and compare responses, because that comparison is the only real proof that a migrated feature behaves the same. And it needs a documented owner, because 'the migration team' will be disbanded before the migration is done. ## The old system will not die on its own The second dishonesty is the belief that once new features go to the new system, the old one will fade. It does not fade. It stays exactly as large as it was, because the features that were never important enough to migrate are still important enough that someone depends on them. The long tail is the migration. The first sixty percent of traffic moves in the first third of the project, because it is the big, well-understood flows. The last ten percent takes longer than everything before it, because it is the monthly reconciliation job, the export nobody documented, the endpoint an external partner calls once a quarter. Each one requires an investigation to find out whether it still matters. The honest plan budgets for this. It lists every capability of the legacy system before the first line of new code, including the ugly ones, and assigns each a decision: migrate, retire or replace with something simpler. 'We will figure it out as we go' is how migrations reach eighty percent and stall forever, with two systems to run instead of one. ## Data is the hard part, not code The third dishonesty is planning the migration around code paths and leaving data for later. Every strangler fig I have seen struggle struggled on data. The new system needs the customer record. The old system owns it. Now either both write to the same tables, which resurrects every problem of a shared database, or the data gets copied, which means a sync job, a conflict policy and a period where the two systems disagree. The honest sequence migrates data ownership as part of each slice, not as a separate phase. When the 'addresses' feature moves, the addresses table moves with it, and the old system starts reading addresses through the new system's API. That is slower per slice and it is the only approach where the old system actually gets smaller. A migration that moves all the code and none of the data has moved nothing. ## What honest looks like on a timeline I ask for four things before endorsing a strangler fig plan. A complete inventory of the legacy system's capabilities, with an explicit decision on each. A facade with metrics, dual-running and a named owner. A data ownership plan per slice. And a definition of done that includes switching the old system off, with a date, and a budget for the long tail that is at least as large as the budget for the first half. With those, the pattern delivers exactly what it promises: continuous delivery of value while the old system shrinks. Without them, it delivers a second system, a proxy nobody owns and a legacy that is now harder to retire than it was on day one, because the migration itself became a dependency. The fig does strangle the tree. But only if someone keeps watering the fig after the first year, when it stops being exciting. --- ## Guardrails that do not ruin the product 2026-07-26 · ai, ux I have used products where the AI feature refuses so often that people learn to route around it, and products where it never refuses and eventually says something that ends up in a screenshot. Both are guardrail failures. The first team put a filter at the end and tuned it for fear. The second team did not put one anywhere. The interesting work is in between: guardrails that shape what the model does so well that the user rarely sees a wall, and when they do, the wall says something useful. ## Constrain the task, not the user The cheapest guardrail is a narrow task. A model asked to "help with anything" needs a huge safety surface. A model asked to "summarize this ticket into three fields" barely has room to misbehave. I design AI features as specific tasks with structured inputs and structured outputs, because a schema is a guardrail: if the output must be one of five categories, the model cannot editorialize. Most of the guardrail budget teams spend on output filters is a symptom of having given the model too much freedom in the first place. ## Layers, in order of cost Input validation first, which is normal engineering: reject empty, oversized or malformed inputs before spending a token. Then the prompt itself, with clear scope and explicit instructions for what to do when the request is out of scope, which should be a structured signal, not a lecture. Then output validation against the schema, with a retry on failure. Then, only for the outputs that reach a person, a classifier or a small model checking for the specific risks that matter in this product. Each layer catches what the previous one missed and costs more than it. Putting the expensive layer first is how you get a slow feature that still leaks. ## Refuse well When a guardrail fires, what the user sees is the product. A generic "I cannot help with that" teaches the user that the feature is unreliable. A specific message that says what the feature can do and offers the closest valid action keeps them inside the product. The model should not write this message. The application should, from a typed reason the guardrail returned: out of scope, missing information, sensitive content, low confidence. Each reason gets its own copy, its own next step, and its own metric, because a spike in one reason is a signal about the product and a spike in another is a signal about abuse. ## Measure the false positives Every guardrail has two error rates and teams only track one. They count the bad outputs that got through. They do not count the good requests that got blocked, which is the number the user feels. I sample blocked requests and review them, the same way I sample passed outputs, and I set a target for both. A guardrail with a false positive rate nobody measures will drift toward blocking everything, because every incident tightens it and nothing ever loosens it. ## The human path For the cases that neither the model nor the rules can settle, there should be a path to a person, with the context attached. Not as a fallback of last resort but as a designed part of the flow. The guardrail's job is to sort: handle the safe majority automatically, block the clearly bad, and route the ambiguous to someone who can decide. A system that only knows how to allow or deny has no place to put ambiguity, so it puts it in one of the two buckets and gets it wrong. --- ## Auditing a system you did not build, without offending the people who did 2026-07-24 · audit, career The technical part of an audit is the easy part. I find the failures; that is a method, and I have written about it. The hard part is what happens after: telling a team of capable, overworked people that the system they built has serious problems, in a way that leads to fixes instead of a defensive meeting. I have done this hundreds of times and I have gotten it wrong enough to have learned some rules. ## Every finding has a context, and the context is not incompetence The first thing I do before writing a single finding is reconstruct why the system is the way it is. Missing timeouts usually mean the dependency was fast when the code was written. Scattered writes usually mean the team grew faster than the architecture. A swallowed error usually means an incident at 2 a.m. was patched by someone who needed to go back to sleep. When I can name the constraint that produced the problem, the finding becomes 'here is what changed since this decision made sense' instead of 'here is what you did wrong'. The first gets fixed. The second gets argued. ## Ask before you assert When I find something that looks wrong, I ask about it before I write it down. 'I noticed the payment webhook is not idempotent, was that a deliberate trade-off?' Sometimes the answer is yes and there is a compensating control I did not see, and I have avoided a false finding. Usually the answer is a pause, and then the engineer explains what they meant to do and never had time for. Now the finding is theirs as much as mine, and they will fix it because they already know it is right. ## Evidence, not opinion Nothing goes in the report without a reproduction. Not 'this could cause duplicates' but 'I sent this webhook twice and here are the two rows'. Not 'this query will be slow' but 'here is the query plan and here is the row count in production'. Evidence takes the disagreement out of the room. Nobody argues with a screenshot of their own database. It also protects the team from me: if I am wrong, the evidence shows it, and I would rather be corrected than believed on authority. ## Rank honestly and say what is fine A report that lists forty findings with equal weight is a report nobody acts on. I rank by what will actually hurt: what loses data, what loses money, what exposes users, then everything else. The top three get the detail, the reproduction and the fix proposal. The rest get a line each. And I say, explicitly and early, what the team did well. Not as flattery, but because it is true, and because a team that hears 'your write boundary is excellent, here is where it leaks' listens differently from one that hears only leaks. ## The report is for the fix, not for the record I write the audit so that the person who has to act on it can do so on Monday morning. That means every finding has a concrete next step, an estimate of effort, and a way to verify it is done. It does not mean a long document that proves how thorough I was. If the team fixes the top three findings and never reads page twelve, the audit worked. My job is not to be right in a document. It is to leave the system better than I found it, with the people who built it still willing to talk to me. My father treated patients for forty years, many of whom could not pay him. From what they tell me, he never made anyone feel small for the state they arrived in. I think about that when I open a codebase that is in bad shape. The people who built it were doing their best inside constraints I did not see. The audit is a diagnosis, and a diagnosis delivered with contempt is a diagnosis nobody follows. --- ## CI that tells you the truth 2026-07-22 · infra, engineering The purpose of continuous integration is to answer one question quickly and honestly: is this change safe to merge? Most pipelines I review answer it slowly and dishonestly. They take twenty minutes, they are green when the code is broken, they are red when nothing is wrong, and the team has learned to click 'rerun' until the color they want appears. At that point the pipeline is a ritual, not a signal. ## The three ways CI lies The first lie is the flaky test. A test that fails one run in twenty, for reasons unrelated to the change, teaches the team that red means nothing. Every flaky test lowers the trust in every real failure. The correct response is not to retry it automatically. It is to quarantine it the same day, fix it or delete it within the week, and treat the flake rate as a metric that must stay at zero. The second lie is the green build on broken code. This happens when the tests do not exercise the thing that changed: the migration that never runs in CI, the environment variable that is set in the pipeline but missing in production, the build step that succeeds because it caches the previous output. The pipeline confirms that yesterday's code still works. The fix is to make CI build the exact artifact that will be deployed, from a clean checkout, and to run the migrations against a real database engine, not a mock. The third lie is the red build that is not about the code. A dependency download timed out. A runner ran out of disk. A third-party API used in an integration test was down. These are infrastructure failures dressed as test failures, and they must be separated: a distinct status, a distinct alert, and never a reason for a developer to look at their diff. ## What a truthful pipeline looks like It runs on every push, and the fast checks finish in under five minutes. Type checking, linting, unit tests, a production build. Anything slower than that runs in a second stage that does not block the developer from continuing, but does block the merge. It builds the deployable artifact once and reuses it for every later step. The image or bundle that passed the tests is the one that ships. If the deploy step rebuilds, the tests were about something else. It runs against real dependencies where the risk is real. A throwaway database container for migrations and integration tests. The actual queue. The actual cache. Mocks are for third parties you do not control, and even then, a contract test against a recorded response is better than a hand-written stub that drifts. It reports in a form a human can act on. Not a wall of logs, but the failing test's name, its assertion, and a link to the line. The difference between a pipeline that gets read and one that gets rerun is whether the failure is legible in ten seconds. ## The rules I enforce - A test that flakes is quarantined the day it flakes, and the quarantine list is reviewed weekly. - The main branch is always deployable; a red main is an incident, fixed or reverted within the hour. - No merge without green, no exceptions, including for the person who owns the repository. - Infrastructure failures in CI are reported separately from test failures and never require a code change to resolve. - The pipeline duration is tracked, and anything over ten minutes for the blocking stage gets an engineering task. ## Why it matters more than it seems A pipeline that tells the truth changes how a team behaves. Deploys become boring because green means safe. Reverts become fast because red means broken. Reviews get shorter because the reviewer trusts the machine to check the mechanical things and spends their attention on design. The cost of a truthful pipeline is a few days of setup and a weekly hour of maintenance. The cost of a lying one is paid on every merge, by every engineer, forever, in the form of 'let me just rerun it'. --- ## Idempotency is a business decision 2026-07-21 · system-design, engineering Engineers talk about idempotency as a technical property: call it twice, get the same result. That is true, and it hides the important part. Deciding what "the same result" means is not a technical question. It is a question about money, about promises to customers, and about who is allowed to be wrong. ## The retry that charges twice The classic case. A client calls "charge the card", the network drops after the charge succeeded, the client retries, the card is charged twice. Everyone agrees this is bad. The fix is an idempotency key: the client sends a unique key with the request, the server stores the result under that key, and a retry with the same key returns the stored result instead of charging again. Simple. Now the questions start, and none of them are technical. ## Questions only the business can answer How long do we remember the key? Twenty-four hours means a retry on day two charges again. Forever means storage and a lookup on every request. A payments team might say seven days. A ticketing system might say until the event happens. What if the same key arrives with a different amount? Is that an error, a new request, or fraud? Most payment APIs reject it. A shopping cart might accept the latest. The answer depends on whether your clients are trusted internal services or third parties. What is the unit of "same"? Same key. Same customer, same product, same minute. Same invoice number. A subscription renewal is idempotent per billing period, not per request, and that rule comes from finance, not from the engineer. And the one people forget: what does the retry see? The original response, byte for byte? Or the current state of the resource, which may have moved on? If the first request created an order and the order has since shipped, does the retry return "created" or "shipped"? Support will have an opinion. ## Where idempotency should live The boundary that matters is the one where an action becomes irreversible or expensive: charging money, sending an email, calling a third-party API with a quota, creating a record another system will react to. Everything before that boundary can be replayed for free. Everything after it needs a key. My rule: every write endpoint that a client can retry accepts an idempotency key, and every consumer of a queue treats each message as if it will be delivered twice, because it will. The key is stored in the same transaction as the side effect. If they live in different stores, there is a window where the side effect happened and the key was not written, and that window is exactly where duplicates come from. ## The cost of pretending Teams sometimes skip idempotency because "our clients do not retry". They do. Mobile networks retry. Browsers retry on connection reset. Load balancers retry on certain errors. Your own queue redelivers. If a request can be sent, it can be sent twice, and the only question is whether you decided what happens. I have audited systems where the duplicate rate was under one in ten thousand and it still mattered, because the duplicates were refunds. A rare bug multiplied by a large enough number is a scheduled event. ## How to have the conversation Bring the product owner a list, not a lecture. For each write in the system: what happens if it runs twice, what that costs, and what we would want to happen instead. Most rows will be "harmless, ignore". A few rows will be "we lose money" or "the customer gets two of something". Those few rows are where the business decides the key, the window, and the conflict rule. Then, and only then, the engineer implements it. Idempotency is not a library you add. It is a set of promises you make, and the business has to sign them. --- ## The hypothesis loop, and where software can shorten it 2026-07-18 · science, ai Every scientific result is the output of a loop. Observe, ask, hypothesize, predict, test, update, repeat. The loop is not a metaphor. It has stages, each with a cost and a latency, and like any loop the total time is dominated by the slowest stage. When I think about where software, including AI, can help science, I do not ask what it can do. I ask which turn of the loop it is shortening, and whether that is the turn that should be shorter. ## The stages and their real cost Observing usually needs instruments and patients and time. Nothing digital changes that much. Asking a good question needs someone who knows the field deeply and has noticed something that does not fit. Forming a hypothesis needs the literature, all of it, including the paper from 1998 that already tried this. Predicting needs a model of what you would see if you were right. Testing needs an experiment or a dataset and the patience to analyze it honestly. Updating needs the humility to accept the result. Look at that list and notice which stages are bounded by physics and which are bounded by reading, searching, formatting and remembering. The second group is where months disappear, and it is the group software is good at. ## Where shortening is safe Finding what has already been tried is a search problem. Turning a vague intuition into a precise, falsifiable statement is a structuring problem. Checking whether a hypothesis contradicts existing evidence is a cross-referencing problem. Writing the analysis code for a well-defined test is a programming problem. Each of these can be made faster by tools without touching the scientific content, because the tool is not producing the knowledge. It is reducing the friction between the person and the knowledge that already exists. A researcher who can go from 'I wonder if' to 'here is the exact claim, here is what would falsify it, and here are the four papers that already touched it' in an afternoon instead of a month is not doing worse science. They are doing the same science with the waiting removed. ## Where shortening is dangerous The test itself should never be shortened by software. Neither should the update. If a tool makes it faster to believe something, that is not acceleration, that is a shortcut through the one stage that produces truth. This is why I am suspicious of anything that generates conclusions. Generating candidates is fine. Generating confidence is not. The rule I use: software may shorten any stage whose output a human can check in less time than it took to produce. A list of relevant papers, checkable. A precisely worded hypothesis, checkable. A statistical result on real data is only checkable by rerunning the analysis, and so the tool must make that rerun trivial, not hide it. ## The human stays at the decision The loop has one stage that must stay entirely human: deciding what the result means and what to do next. Not because machines cannot compute it, but because someone has to be accountable for it. In medicine that accountability eventually reaches a patient. A system that shortens the loop while blurring who decided is faster and worse. This is the frame I bring to ELUCENIA: every hypothesis needs evidence, every advance needs scientific validation, every decision needs an accountable human. Shorten the reading. Shorten the searching. Shorten the plumbing. Leave the thinking and the testing exactly as slow as truth requires. --- ## React state that belongs to the server 2026-07-15 · fullstack, nextjs Open a mature React codebase and count the useState calls. Then count how many of them hold something that came from the server: a user, a list of orders, a product's price. In most projects I audit, more than half of client state is a copy of a database row, and every one of those copies is a chance to be wrong. The App Router changed what the right answer is. Before it, the client had to hold server data because there was nowhere else to render it. Now there is, and the question for every piece of state is simply: who owns this? ## Three kinds of state Server state is data the database owns. Orders, users, prices, permissions. The server is the source of truth and the client only ever holds a snapshot. Client state is data that exists only in the browser session: whether a menu is open, which tab is selected, what the user typed but has not submitted. URL state is data that should survive a refresh and be shareable: filters, pagination, the selected date range. The bugs come from putting one kind in another's container. Server state in useState goes stale the moment someone else changes the row. Client state on the server means a round trip to open a dropdown. Filters in useState mean the back button breaks and the link cannot be shared. ## Server state lives in Server Components A Server Component reads the data it needs, at request time, from the database. It passes the fields the client needs as props. There is no loading flag, no error state, no effect that fetches after mount, no cache key to invent. The snapshot is as fresh as the request. When the user changes something, a Server Action performs the mutation and calls `revalidatePath` or `revalidateTag`. The framework re-renders the affected Server Components and streams the new tree to the client. The component that displayed the old value never held it in state, so there is nothing to synchronize. For the gap between click and response, `useOptimistic` holds a temporary value that the server's answer replaces. That is the one legitimate client copy of server state, and it is scoped to a single pending action rather than living in a store. The pattern I remove most often in audits is the fetch in useEffect that sets state. It creates a request waterfall, a flash of empty content, a race when the component remounts, and a stale copy that outlives its usefulness. Moving that fetch into the Server Component deletes all four problems and about thirty lines. ## URL state lives in the URL Search params are the right home for anything a user would want to bookmark or share. A Server Component reads them as a prop and queries accordingly. A Client Component updates them through the router, and the server re-renders. The state is visible, survives reloads, and the back button works. The rule I use is that if two users looking at the same URL should see the same thing, the state that decides it belongs in the URL. If the state is private to this session, such as a half-typed search, it stays in the client. ## When client fetching is still right None of this makes client-side data libraries obsolete. Highly interactive screens that refetch on focus, poll for updates, or need to work offline still benefit from a client cache with its own lifecycle. A real-time dashboard, a collaborative editor, an app that must function on a train with no signal. The mistake is using that machinery by default, for a product page that changes once a day, because it was the only option in an older architecture. Start with the server owning server state. Reach for a client cache when the screen has a requirement the server cannot meet, and be able to say what that requirement is. State has an owner. Put it where the owner is, and most of the synchronization code you have been writing turns out to have been a workaround for putting it somewhere else. --- ## Hallucination is a system property, not a model bug 2026-07-11 · ai, reliability When a feature confidently states something false, the post-mortem usually ends with one word: hallucination. As if the model had a defect that a better model will fix. I read those post-mortems differently. The model did what a language model does: it produced the most plausible continuation of the context it was given. The system asked it a question the context could not answer, gave it no way to say so, and shipped the result to a user with no check in between. The false statement is a property of that whole arrangement, not of the weights. ## The model is doing its job A language model is a plausibility engine. Given a context, it produces text that looks like it belongs there. When the context contains the answer, plausible and true coincide. When it does not, the model still produces plausible text, because that is the only thing it can do. Calling this a bug is like calling it a bug when a search engine returns results for a query with no good matches. The engine ranked what it had. Whether the top result should be shown as the answer is the caller's decision. ## Where the system invites hallucination I can usually predict where a feature will hallucinate by reading its architecture. Retrieval that returns nothing relevant, followed by a prompt that says "answer the question" with no permission to decline. Contexts truncated silently so the model never sees the sentence that mattered. Questions about numbers, dates and names, which are exactly the things a model is worst at generating and best at extracting. Open-ended prompts where a structured one would have forced the model to cite or abstain. Each of those is a design choice, and each one makes false output more likely, independent of which model sits underneath. ## Design that reduces it The countermeasures are structural, not clever. Give the model the facts and tell it to use only those. Give it an explicit way out: a field in the structured output that says "not found in context", which the interface renders honestly. Verify the extractable parts in code: if the model cites a passage, check that the passage exists and contains the claim. Separate generation from decision: the model drafts, code validates, a person approves where the stakes require it. And measure the rate: sample production outputs, label them, track the fraction that is unsupported. A rate you measure goes down. A rate you name and never measure stays where it is. ## Measure it as a rate The word hallucination hides a number, and the number is what matters. In a feature with a good retrieval layer, an explicit abstain path and validation of cited facts, the unsupported-claim rate can be very low, and low enough for the product. In a feature with none of those, it can be high enough that the model choice barely matters. When I audit, I ask for the rate and for the eval set that produced it. Teams that have the number are already treating the problem as a system property. Teams that have the word are still waiting for a better model. ## Accountability This matters most where the output informs a real decision. In work that touches health, or science, or money, every claim needs evidence behind it, every result needs validation, and every decision needs a person whose name is on it. That is the principle behind ELUCENIA, and it is not a constraint on the model. It is the architecture that makes the model safe to use. Hallucination is what happens when a system skips one of those steps and hopes plausibility is enough. Sometimes it is. In the places that matter, it never is. --- ## Why most microservices should have been modules 2026-07-08 · system-design, architecture I have reviewed enough systems to say this plainly: most microservices I meet are modules that were given a network address too early. The team paid the full price of distribution and got almost none of the benefits. ## What a network boundary actually costs When a function call becomes an HTTP or gRPC call, five things change at once. Latency: an in-process call is nanoseconds; a call across the network is a millisecond on a good day, and a request that touches six services pays it six times, serially. Partial failure: a function either returns or throws. A remote call can time out after doing the work, leaving you unsure whether it happened. Versioning: you can no longer change a signature and fix all callers in one commit; the old and new shapes must coexist while deploys roll. Transactions: an order and its inventory reservation used to commit together. Now they live in two databases and you need sagas, compensation and outbox tables to get back something weaker than what you had. Observability: a stack trace becomes a distributed trace, which you have to build and maintain. None of these are exotic problems. They are all solvable. But each one is a permanent tax, paid on every feature, forever. ## The modular monolith The alternative is not a big ball of mud. It is a single deployable with hard internal boundaries: one module per domain, each with its own public interface, its own tables that no other module reads directly, and an import rule enforced by tooling so that the boundary cannot be crossed casually. You get the thing microservices promised, clear ownership and independent evolution, while keeping one deploy, one database transaction, one trace, one stack trace. In TypeScript this is folders plus lint rules and a package-level index that is the only allowed import. Simple, and it holds up if the team is disciplined about it. ## When a service earns its boundary A service should exist only when a boundary pays for itself, and there are four honest reasons. - Independent scaling: one part of the system needs 50 instances and the rest need 2, and they compete for the same resources. - A different runtime: the video transcoder needs a GPU and Python; the API is Node. - Team ownership: two teams ship on different cadences and keep blocking each other's deploys. - Security or compliance isolation: the part that touches card data or health records must live in a smaller, separately audited perimeter. If you cannot name one of those, you are paying the tax for a diagram. ## Signals a service should have been a module Some patterns show up over and over. Two services that are always deployed together, in a fixed order, because they cannot work with mismatched versions. A service whose only client is one other service. A "shared" database that three services write to. Endpoints that exist only so service A can read service B's tables. A feature that needs a pull request in four repositories. Any of these means the boundary is in the wrong place, or should not exist. ## Extraction is cheaper than merging The good news is that the mistake is asymmetric. Extracting a well-bounded module into a service later is a mechanical job: the interface already exists, the data is already separate, you swap a function call for a client. Merging two services that grew apart, with duplicated models, divergent validation and two sets of migrations, is months of careful work that nobody wants to fund. So the default should be a module. Design it as if it might become a service, keep the boundary clean, and let it stay in the monolith until one of the four reasons shows up. Most never will, and that is the point. --- ## Say no to the rewrite 2026-07-06 · career, architecture Every engineer eventually stands in front of a codebase and feels the pull. It is old, it is inconsistent, nobody understands the middle third of it, and the thought forms clearly: we should rewrite this. I have felt it many times. In 600+ projects I have watched the rewrite proposed, approved, started, and, in most cases, quietly abandoned or delivered so late that it solved a problem the business no longer had. This is the case for saying no. Not always. But by default, and with the burden of proof on the rewrite. ## What the rewrite promises and what it delivers The rewrite promises a clean slate. What it delivers is a second system that must reach feature parity with the first before anyone can use it, while the first keeps changing because the business does not pause. You are chasing a moving target with a team that is now split between maintaining the old and building the new. The part nobody prices in is the knowledge locked in the old code. That strange conditional in the billing module is not a mistake. It is a rule someone learned from a customer complaint in 2019 and nobody wrote down. The rewrite deletes it. Six months after launch, the complaint comes back. ## The questions I ask before approving one When a team brings me a rewrite proposal, I ask them to answer these in writing. - What specifically cannot be done in the current system? Not "it is hard to maintain". Which change, requested by whom, is blocked? - Have you tried to fix that one thing in place, and what happened? - What is the plan for the period when both systems exist? Who owns the old one? - If the rewrite stopped at 60%, what would you have? - What are you going to do differently so that the new system does not become the old one in five years? Most proposals do not survive the first question. The honest answer is usually "nothing is blocked, we just do not like it". That is a real cost, but it is not a rewrite-sized cost. ## What to do instead The alternative is not "live with it". It is to treat the codebase like a city, not a building. You do not demolish a city because one district is run down. You fix the district. Find the module that hurts the most. Put a boundary around it: a clear interface, tests at the edge, a contract that the rest of the system talks through. Then replace what is inside the boundary, and only that. When it works, pick the next module. This is slower per module and faster overall, because the system keeps running and the business keeps shipping the whole time. The other thing to do is to write down the knowledge before it is lost. Every strange conditional gets a comment explaining the customer complaint behind it. Every undocumented rule becomes documented. This alone reduces the pull toward the rewrite, because half of that pull is fear of what you do not understand. ## When the answer is yes There are rewrites I have approved. The platform was being discontinued. The language had no maintained runtime. The data model was wrong at the root and every feature was a workaround for the same wrong assumption. In those cases the answer to "what cannot be done" is specific and severe, and the rewrite is not a preference, it is a deadline. Even then, I insist on the boundary approach. The new system takes over one responsibility at a time. The old one is retired module by module, not on a launch day. ## The career part Saying no to the rewrite is unpopular. The team wants it. The new engineer wants it. The rewrite is exciting and the fix is boring. But the engineers I trust most are the ones who can look at an ugly, working system and say: this is ugly, it works, and here is the smallest change that makes it less ugly without stopping it from working. That sentence, repeated for years, is how you build a reputation for being the person who ships. --- ## Containers are not isolation 2026-07-03 · infra, security Somewhere along the way, 'it runs in a container' became shorthand for 'it is isolated'. I hear it in design reviews when someone proposes running untrusted code, third-party plugins, or a customer's uploaded script. The container is offered as the safety net. It is not one. Or rather, it is a net with holes that are well documented and easy to forget. ## What a container actually is A container is a regular process on the host, started with a few kernel features that restrict what it can see. Namespaces give it its own view of the process table, the network interfaces, the filesystem mounts, the hostname. Control groups limit how much CPU and memory it can consume. A restricted set of capabilities removes some of the things root can normally do. A seccomp profile filters which system calls it may make. That is the whole trick. The process still shares the host kernel with every other container. Every system call it makes is handled by the same kernel that handles the database next door. A kernel bug reachable from inside a container is a kernel bug reachable from the host. Virtual machines put a hypervisor between the guest and the hardware. Containers put a table of rules between the process and the kernel. Those are different amounts of distance. ## The holes that are not bugs The dangerous part is that most container escapes I see in audits are not exploits. They are configuration. A container started with the privileged flag has essentially all of root's capabilities on the host, and people use that flag to make something work without reading what it does. A container that mounts the container runtime's socket can start new containers with any configuration it likes, including privileged ones. A container that runs as root inside, on a host where user namespaces are not remapped, is root on any file it can reach through a mount. Then there are the softer holes. A container with no memory limit can starve the host. A container with no PID limit can fork-bomb the node. A container that can reach the cloud provider's metadata endpoint can often fetch credentials for the instance role, which is frequently far more powerful than the application needs. None of these require escaping the container. They require the container being allowed to do things it never needed to do. ## What I actually check - The process inside runs as a non-root user, and the image is built so that it cannot become root. - The root filesystem is mounted read-only, with explicit writable volumes only where needed. - All capabilities are dropped and only the specific ones required are added back, which is usually none. - CPU, memory and PID limits are set, and the memory limit is tested against real workload peaks. - No host paths, no runtime socket, no host network mode, no privileged flag, and the metadata endpoint is blocked unless the workload genuinely needs it. This list is not exotic. It is the default recommendation in every hardening guide. It is also violated in most systems I review, usually because the base image runs as root and nobody changed it, or because a build step needed a capability once and it was never removed. ## When you actually need isolation If the workload is your own code, hardened containers on managed infrastructure are a reasonable boundary. The threat model is a compromised dependency or a bug in your service, and the hardening above contains most of what such a compromise can do. If the workload is code you did not write and do not trust, meaning customer scripts, plugins, generated code, or anything from a multi-tenant sandbox feature, a container is not enough. That is where you reach for a real boundary: a microVM per workload, a sandboxed runtime designed for untrusted code, or a separate host per tenant. These cost more and start slower. That is the price of the actual isolation the word 'container' implied all along. The rule I give teams is simple. Containers are a packaging and scheduling tool with useful resource limits. They are not a security boundary between distrusting parties. Design accordingly, and you will not be the one explaining how a customer's upload read another customer's environment. --- ## Forms are distributed systems 2026-07-01 · fullstack, engineering A form looks like the simplest thing in web development. A few inputs, a button, a handler. I have found more money-losing bugs in forms than in any queue, cache or database I have ever audited. The reason is that a form is a distributed system in disguise. There are two machines, the browser and the server. There is a network between them that drops packets, times out and delivers late. And there is a user who clicks twice, closes the tab mid-request and opens the page in two windows. Every classic distributed systems failure is available, and most forms are written as if none of them exist. ## Duplicate submission is the default The first failure is the double submit. The user clicks, nothing visibly happens for 800 milliseconds, they click again. Two requests arrive. Two orders are created, or two payments are charged, or two emails go out. Disabling the button on click is not a fix. It is a mitigation that fails the moment the network retries a POST, the browser resends after a back navigation, or a second tab submits the same draft. The fix lives on the server: an idempotency key generated when the form is rendered, sent with the submission, and checked before any side effect. The second request with the same key returns the result of the first. In Next.js with Server Actions, the key travels as a hidden field or as an argument. The action looks it up in the database in the same transaction that creates the record. If it exists, return early. That one table with a unique constraint has saved more revenue than any amount of frontend debouncing. ## Validation happens twice, and that is fine Client-side validation is for the user. Server-side validation is for the system. They are not redundant, they have different customers, and skipping the second one because the first exists is the security hole I find most often in audits. The trick to keeping them in sync is one schema, shared. Define the shape once with a runtime validation library, import it in the Client Component for instant feedback and in the Server Action for the real check. When a field is added, both sides learn about it from the same file. Error reporting has to be structured. A form that returns 'something went wrong' has thrown away the information the user needs to fix it. I return a map of field name to message, and the framework's form state hook puts each message next to its input. The useActionState hook exists precisely for this: it carries the previous result back to the client without a client state library. ## State that survives the network The user typed for four minutes and the submission failed because their session expired. If the form resets, you have lost a customer. Progressive enhancement is the baseline: a form that posts to a Server Action works before JavaScript loads and keeps the typed values on the server response. For long forms I persist drafts, either to local storage or to the server as an explicit draft record, keyed to the user. Autosave every few seconds, restore on load, and tell the user it happened. That is a small feature that turns an abandonment into a return visit. Optimistic updates are the same idea in reverse. Show the result immediately, send the request, and reconcile when the server answers. The useOptimistic hook handles the mechanics, but the design decision is yours: what does the user see if the server says no? Decide that before you ship it. ## Concurrency and the second tab Two windows edit the same record. Both submit. Last write wins, and the first editor's work is silently gone. This is a classic lost update, and the defense is a version number. The form carries the version it loaded. The update checks that the version in the database still matches. If not, reject with a message that says someone else changed this, and show the diff. That is optimistic concurrency control, and it is one integer column. I add it to every table a human edits. None of this is exotic. It is the same idempotency, validation, persistence and versioning that any distributed system needs. The form is just the place where most teams forget they are building one. --- ## Core Web Vitals are an architecture problem 2026-06-30 · fullstack, performance When a team asks me to help with Core Web Vitals, they expect a performance engineer. They get an architect. The scores are a symptom, and the disease is almost always structural: where the data comes from, where the boundary between server and client sits, and what the page is allowed to do before it knows what it is showing. Three metrics matter. Largest Contentful Paint, how long until the main content is visible. Interaction to Next Paint, how long the page takes to respond when the user does something. Cumulative Layout Shift, how much the page jumps around. Each one maps to an architectural decision. ## LCP is a data flow decision The largest element on most pages is a hero image, a headline or a product card. It cannot paint until the server has the data to render it, and the server cannot have the data until the slowest query in the critical path returns. So the question is not how to optimize the image. It is why the image is waiting for a query that has nothing to do with it. With Server Components and Suspense, the page shell and the hero can render immediately while the recommendations panel streams in later. That is an architectural choice about what is critical, made per page, and it moves LCP more than any image format ever will. The second LCP killer is the request waterfall. A layout fetches the user, then a page fetches the account, then a component fetches the plan, each waiting for the previous one. Start every independent fetch at the top and await them together. The waterfall is a code structure problem, and it shows up as seconds. Then the boring parts that are still architecture: the image is served with explicit dimensions and a priority hint, the fonts are self-hosted and preloaded, the CDN is actually caching the static assets and the HTML has a sane Cache-Control. I check the headers before I check the code. ## INP is a boundary decision Interaction to Next Paint is bad when the main thread is busy. The main thread is busy when there is too much JavaScript, and there is too much JavaScript when the server-client boundary is drawn too high. A layout marked with `'use client'` for a menu toggle drags the entire subtree into the bundle, and every one of those components hydrates before the page can respond to a click. The fix is to move the boundary down. Interactive islands, small and specific, with Server Components around them. Third-party scripts are the other offender: a chat widget, an analytics tag and a consent banner can each consume more main thread time than the application itself. Load them after interaction, or not at all, and measure each one alone before deciding. Long tasks inside event handlers are the last piece. If a click triggers a synchronous computation over ten thousand rows, no framework will save you. Move the work to the server, a worker or a transition, and paint first. ## CLS is a contract decision Layout shift happens when the page does not know how big something is going to be. An image without dimensions. A banner injected after load. A font that swaps to a different metric. Each one is the browser being asked to guess. The architectural answer is that every asynchronous element reserves its space before it arrives. Skeletons with the final dimensions. Aspect ratios on media. Fallback fonts with adjusted metrics. A Suspense boundary whose fallback is the same height as its content. That is a contract between the server and the layout, and it is designed, not tuned. ## Measure in the field, not the lab Lab scores from a desktop on fiber are a lie. The users who fail the thresholds are on a mid-range phone on a congested network, and only real-user monitoring shows them. Ship the web vitals library, send the numbers to your own backend with the route name attached, and look at the 75th percentile per route, because that is what the thresholds are defined on. When you do, you will find the same thing I find in audits: the slow route is slow because of a decision someone made about data, not about pixels. Fix the decision, and the score follows. --- ## Why a software engineer cares about science 2026-06-28 · science, career People ask me why a full-stack engineer who spends his days in Next.js and TypeScript writes about laboratories, hypotheses and cardiovascular data. The honest answer is that I never saw a difference. Science is a system. It takes noisy observations as input, runs them through a process with rules, and produces knowledge that is supposed to survive contact with reality. That is exactly what I do for a living, with worse inputs and shorter deadlines. ## The method is a design pattern Every engineer already practices a degraded version of the scientific method. You see a bug, you form a guess about the cause, you design a change that would prove or disprove the guess, you run it, you look at the result. The difference between a junior and a senior engineer is mostly how honest that loop is. Juniors change three things at once and declare victory when the error disappears. Seniors change one thing, predict the outcome before running it, and get suspicious when the prediction is wrong in the right direction. Science formalized that discipline centuries before we had version control. Controls, blinding, pre-registration, replication: each one is a countermeasure against a specific way humans fool themselves. When I audit a system and find the failure nobody else found, it is rarely because I know more. It is because I refused to accept an explanation that had not been tested. ## I studied three years of Biomedicine and left Before software I spent three years in a Biomedicine course. I did not finish. I was better at reading a protocol and asking why step four came before step five than I was at executing it. What stayed with me was the shape of the work: the slow, careful accumulation of evidence, the respect for what you do not yet know, and the frustration of watching good questions die because the tooling around them was terrible. That frustration never left. Across 600+ projects I have watched businesses get better software every year while researchers still email spreadsheets to each other and rebuild the same pipeline in every lab. ## My father was a doctor, and I barely knew him My father, Pedro Moretti Guedes, practiced medicine in São Paulo for more than forty years and treated for free the people who could not pay. I met him once, as a teenager, for about two hours. Almost everything I know about him comes from his patients. That is a strange way to inherit a vocation, through other people's memories, but it is the one I got. I did not become a doctor. I became someone who builds systems, and at some point I realized the two are not as far apart as they look. A clinic is a system. A diagnosis is a hypothesis with a feedback loop. The person who cares about a patient and the person who cares whether a pipeline is reproducible are trying to protect the same thing: a decision that is going to affect a life. ## What an engineer can actually contribute I am not going to discover a new mechanism of heart disease. Clinicians and scientists will. What I can do is what I have always done: look at the system around them, find where it silently loses information, and build the boring, reliable parts that make the interesting parts possible. Provenance. Reproducibility. Validation that is not a phase at the end. Interfaces that make the honest path the easy path. That is why ELUCENIA exists, and why it is global from the start: a medical and scientific network across countries and specialties. Its principles are not complicated: every hypothesis needs evidence, every advance needs scientific validation, every decision needs an accountable human. Those are engineering principles as much as scientific ones. I just happened to learn them from both sides. Caring about science is not a hobby for me. It is the same job, pointed at the questions that matter most. --- ## AI pair programming changed what I put on the whiteboard 2026-06-26 · ai, system-design For most of my career the whiteboard was where I worked out what to type. Boxes were modules, arrows were calls, and the session ended when I could see the code in my head. Since a model started writing most of the code with me, that kind of drawing has become nearly worthless. The model can produce the module from a sentence. What it cannot produce is the sentence. So the whiteboard moved up a level, and honestly it moved to where it should have been all along. ## Less code shape, more failure shape The first thing that left the board was structure for its own sake. I no longer sketch class hierarchies or folder layouts, because the model will propose a reasonable one and I will accept or adjust it in minutes. What replaced it is failure. For every arrow I draw, I now write what happens when the thing at the other end is slow, absent, duplicated or lying. That used to be a second pass I did if there was time. Now it is the first pass, because the failure behavior is the part the model will get wrong if I do not specify it, and it will get it wrong confidently. ## Invariants and boundaries The second thing that grew on the board is a list, usually in a corner, of statements that must always be true. An order is never charged twice. A tenant never sees another tenant's rows. A published result always has a source. Those sentences are the real specification of the system, and they translate directly into prompts, into tests and into review checklists. When I hand a task to the model, the invariants go with it, and when I review the output, I check each one. The boxes and arrows are scaffolding. The invariants are the building. ## Data flow with ownership The third change is that every piece of data on the board now has an owner and a lifecycle, drawn explicitly. Where it is created, who can change it, when it is deleted, where it is copied and whether the copies can drift. Models are very good at writing code that moves data around, and very bad at noticing that the data now exists in two places with two truths. That failure, the one nobody finds until the audit, is a whiteboard failure. If ownership is drawn, the code follows. If it is not drawn, the code will be written without it, quickly and neatly. ## What the whiteboard is now for The whiteboard is now where I decide the things the model cannot decide: which trade-offs the product accepts, which consistency the data needs, what the user should see when something breaks, what the cost ceiling is, who signs off on the irreversible actions. Those decisions used to be spread thin across a hundred small coding choices, made half-consciously while typing. Now they have to be made deliberately, in advance, in words, because the typing happens elsewhere and it happens too fast to carry judgment along with it. ## The meeting changed too A side effect I did not expect is that whiteboard sessions became more useful for the people who do not write code. When the board is full of module names, a product manager or a domain expert nods politely and waits. When the board is full of failures, invariants and ownership, they have opinions, and their opinions are usually right, because those are questions about the business, not about syntax. The model took the syntax. What is left is the part that was always the actual work, and it turns out more people can do it than we thought. --- ## Domain events versus notifications: they are not the same thing 2026-06-24 · architecture Most 'event-driven' systems I audit are not event-driven. They are notification-driven with an event bus in the middle. The difference sounds academic until you try to add a second consumer, replay history or explain to a product manager why the order count in two dashboards does not match. A domain event is a fact: something happened in the business, it is named in the past tense, and it carries enough information for a stranger to understand what changed. OrderPaid, with the order id, the amount, the currency, the payment method and the timestamp. A notification is a poke: 'something about order 4471 changed, go look'. It carries an id and maybe a type, and expects the consumer to call back and fetch the state. Both are legitimate. They solve different problems, and every mixed-up system I have seen picked one while needing the other. ## What a domain event promises A domain event promises three things. It is immutable, because it describes the past. It is self-contained, so a consumer can act on it without calling the producer back. And it is meaningful in the language of the business, not the language of the database. 'CustomerAddressChanged' is a domain event. 'customers table row updated' is not, even if it travels through the same broker. Those promises make domain events powerful. You can replay them to rebuild a read model. You can add a consumer a year later and it will understand history. You can audit what happened without joining logs across three services. Analytics, notifications, fraud checks, fulfillment: each subscribes and does its own thing, and the producer never learns they exist. The cost is that the producer must decide, at emit time, what is worth recording. The payload is a contract. Change it carelessly and every consumer that parsed the old shape breaks, sometimes silently, months later, when a replay hits an old version. ## What a notification promises A notification promises almost nothing, and that is its strength. It says 'go check'. The consumer calls the producer's API and gets the current state, so it always sees the truth as of now, not as of the event. There is no payload contract to maintain beyond the id. The producer can change its internal model freely. The cost is coupling in the other direction. The consumer depends on the producer being available and fast when the notification arrives. Ten consumers means ten callbacks per change. Replaying notifications is pointless, because the state they point at has moved on. And there is no history: you know that order 4471 changed twelve times, but not what changed each time. ## Where the confusion does damage The most common failure is emitting notifications and treating them as events. A team publishes 'OrderUpdated' with just an id. Then someone builds a read model by consuming it and fetching the order. Then someone else adds a fraud check on the same topic. Then the orders API has a bad afternoon and every consumer's callbacks fail, the read model drifts, the fraud check misses a window and nobody can rebuild anything because the topic contains no facts, only pointers to state that has since changed. The opposite failure is quieter. A team emits rich domain events, then a consumer ignores the payload and calls the API anyway 'to be safe'. Now you pay the cost of both: a payload contract to maintain and a synchronous dependency on the producer. The event bus becomes a very expensive way to trigger HTTP calls. The third failure is naming. 'OrderChanged' is not a domain event, because the business does not think in 'changed'. It thinks in placed, paid, packed, shipped, refunded. If an event name would not appear in a conversation between two people who work in that domain, it is a technical notification wearing a domain costume. ## How I decide I ask what the consumer will do. If it needs to react to a specific business fact and it should keep working when the producer is down, it needs a domain event with a real payload. If it just needs to refresh a cache or a screen and it is fine to fetch the latest state, a notification is cheaper and safer, because there is no contract to evolve. Then I ask about history. If anyone will ever want to know what happened, in order, with details, only domain events give you that. If nobody will, do not pay for it. And I name them honestly. Facts in the past tense, with the vocabulary of the people who run the business. Pokes as pokes: 'OrderRefreshRequested', with an id, and no pretense. A system where the two are visibly different is a system where a new engineer can tell, from the name alone, what they are allowed to rely on. --- ## Faith, family and uptime 2026-06-21 · life I keep systems running for people, and I have a family and a faith that both matter more to me than any system. For a long time I treated those as separate lives, one of which paid for the other. I do not anymore. This is about how they connect, in practical terms, without pretending the connection is tidy. ## Uptime is a promise to people you will never meet When a system I designed goes down, someone somewhere cannot do a thing that matters to them: pay a bill, book an appointment, submit a form on the last day. I never meet them. Uptime is a promise to strangers. My father was a doctor in Campo Limpo for more than forty years and treated for free people who could not pay. I met him once. But that fact about him, that the promise was to people who could give nothing back, has shaped how I think about reliability. The system should work for the user who has no recourse, no support contract, no way to complain. That is not a business rule. It is closer to a moral one, and my faith is where I get it from. ## Family sets the shape of the work Natália and I have built a life in a small city in Paraná, and the shape of my working day is set by that life, not the reverse. I do not do heroic on-call. I do not take the contract that requires me to be reachable at all hours, because the cost of that lands on people I love rather than on the client. This has made me a better engineer, not a worse one. If I cannot be woken at 3 a.m., the system has to be designed so that it does not need me at 3 a.m. That means alerts that are actionable, runbooks that someone else can follow, failure modes designed in advance, and a strong preference for boring technology that fails in known ways. The constraint forced the discipline. Family is, in a real sense, the reason my systems are reliable. ## Faith changes what I do on the bad days I have written about 2020, when I lost every contract in a few weeks. What faith did for me that year was not dramatic. It gave me a reason to do the next small step on days when the outcome was not visible. That is also, exactly, what you need on the second day of an incident, when the fix from yesterday did not hold and the client is losing patience. The temptation on a bad day is to panic-deploy, to change several things at once, to make the number move. Faith, for me, is what lets me be calm enough to change one thing, measure, and change the next. I am not saying you need faith to do that. I am saying it is where I get it. ## The order matters I hold these in an order: faith, family, work. Not because work is unimportant but because when the order is inverted, everything degrades. The engineer whose identity is the system will burn out or become impossible to work with. The one who knows the system is third on the list can walk away from it at night, and comes back the next morning with the perspective needed to see the failure that was invisible at midnight. ## What is usable here Whatever you hold above your work, put it above your work explicitly, and let the constraints it creates shape the engineering. If you cannot be woken, design for not being woken. If you serve people who cannot pay, design for the user with no recourse. If you need calm on the bad day, find where you get calm and go there before the bad day comes. The systems will be better for it. Mine are. And the research project I am building now, a global medical and scientific network to accelerate discovery, comes from the same order: it is an attempt to be useful to people I will never meet, the way my father was to his patients in Campo Limpo, using the only tools I have. --- ## Evaluating LLM features like an engineer, not a fan 2026-06-19 · ai, audit The most common state of an LLM feature I get called to audit is this: it shipped, people like it, and nobody can tell me whether last week's prompt change made it better or worse. There is no dataset, no score, no baseline. There is a demo that impressed a stakeholder and a few screenshots in a channel. That is fandom, not engineering. Engineering is when you can be wrong in a measurable way. ## Vibes are not a metric A model output looks good on the first read almost by construction. Fluent text is persuasive. The failures are in the second read, in the rare input, in the case that was not in the demo. The only defense is to stop reading individual outputs as a judge of quality and start reading aggregate results over a fixed set of inputs. Once you have that, every change to the prompt, the retrieval, the model or the schema becomes a number that went up or down, and the conversation changes from taste to evidence. ## Build the dataset before the feature I build the eval set the way I write tests for a state machine: before the implementation, from the spec. Fifty to a few hundred inputs is enough to start. They must include the boring middle, the edge cases that matter to the business, the inputs that should be refused, and a handful that are actively hostile. For each, I record either an expected output, a set of required facts, or a grading rubric. Real production traffic, anonymized and labeled, is the best source, and the set grows every time a user reports a bad answer. Every bug becomes a case, the same as any other regression suite. ## Three kinds of evals Some checks are deterministic: the output parses against the schema, the date matches the source, the forbidden phrase is absent, the cited passage exists. Those are cheap and run on every commit. Some are similarity-based: the answer is close enough to a reference, measured by a tolerant comparison. And some need judgment: is this summary faithful, is this reply appropriate. For those I use a model as a grader, with a rubric, and I treat that grader as a component with its own tests. The mix matters. A feature evaluated only by a model judge has a blind spot exactly where the judge and the generator share assumptions. ## Judges need judging Before I trust a model grader, I have humans label a sample and I measure agreement. If the grader agrees with people less often than two people agree with each other, it is not a grader, it is noise with a confident tone. I also check for the obvious biases: preferring longer outputs, preferring its own style, preferring the first of two candidates. Calibrating the judge takes an afternoon and it is the afternoon that makes every subsequent number meaningful. ## Evals as a regression suite Once the set exists, it runs in CI like any other test, with a threshold and a report. A prompt change that drops faithfulness by a few points blocks the merge, the same as a failing unit test. A new model version is evaluated before it is enabled, not after users complain. And the report is a table anyone can read: cases, scores, diffs. That is what turns an AI feature from something a team is proud of into something a team is responsible for. The fans get to keep their enthusiasm. The engineers get to keep their jobs. --- ## Rate limiting without hurting the users you want 2026-06-17 · system-design Most rate limiters I audit were written in an afternoon, after an incident, by someone who was angry at a bot. They work. They also quietly throw 429s at the company's best customers, and nobody notices because the best customers rarely complain. They just retry, or leave. A rate limiter has one job: protect the system from abuse and from accidents. It is not a punishment mechanism. If a paying user hits your limit during normal use, the limit is wrong, not the user. ## Pick the algorithm for the traffic you actually have Two algorithms cover almost everything. A token bucket gives each client a bucket that refills at a fixed rate, say 10 tokens per second, up to a capacity of, say, 50. Each request spends a token. That capacity is your burst allowance: a client can fire 50 requests at once after being idle, then settles into 10 per second. Real users are bursty. A page load fires a dozen calls, then nothing for a minute. Token bucket models that well. A sliding window counts requests in the last N seconds and rejects when the count passes the limit. It is stricter and easier to explain in documentation ("100 requests per minute"), but it has no natural notion of burst, so a user who reloads a dashboard three times can get blocked for something completely reasonable. My default: token bucket for user-facing APIs, sliding window for things that must be strictly bounded, like password attempts or SMS sends. Avoid fixed windows that reset at the top of the minute; they allow double the limit at the boundary and clients learn to exploit that. ## Limit by identity, not by IP IP-based limits are the most common mistake I see. An entire office, a university campus, a hospital, or a mobile carrier's NAT can share one public IP. Limit by IP and you block a thousand legitimate people because one of them wrote a loop. Limit by API key, user id, or session when you have one. Use IP only for unauthenticated endpoints, and make those limits generous, because you have no idea who is behind the address. If you must combine the two, make IP the wide net and identity the fine one. ## Not every endpoint deserves the same limit A global limit per user is a blunt tool. Reads and writes have different costs and different abuse profiles. Reading a profile is cheap and idempotent; creating an order touches inventory, payments and email. Give writes a separate, tighter budget. Then go further and protect expensive endpoints specifically. Search with wildcards, report exports, PDF generation, anything that fans out to several services. These are where an accidental client loop takes you down, and they need their own bucket, sized by cost, not by request count. Some teams assign a cost per endpoint and charge it against a single budget; that works well when the costs are honest. Tier the limits per plan. A free tier, a paid tier and an enterprise tier should not share numbers, and the plan should be readable from the same identity you limit by. This turns rate limiting from a defensive wall into a product feature. ## Reject in a way clients can cooperate with Return 429, not 403 or 500. Include a Retry-After header with seconds, and ideally the standard RateLimit headers so well-behaved clients can slow down before hitting the wall. A client that knows when to retry stops hammering you; a client that gets an opaque error retries immediately, in a loop, and makes the incident worse. Never rate limit health checks, and be careful with webhooks from partners you asked to send you traffic. ## Observe before you enforce The most important advice here: run the limiter in shadow mode first. Compute the decision, log it, count it, but do not reject. After a week you know exactly who would have been blocked and why. Almost every time, the list contains your biggest integration partner, an internal dashboard, and a mobile app version with a retry bug. Fix those, tune the numbers, then flip enforcement on. Keep the shadow metrics running after that. When someone's limit needs to move, you want the data before the support ticket. --- ## From papers to questions: what an instrument for research should do 2026-06-14 · science, ai Most tools that promise to help researchers with the literature do one thing: they make papers shorter. Summaries, abstracts of abstracts, chat over PDFs. That is useful the way a faster elevator is useful. It does not change what you do when you get to the floor. When I think about what an instrument for research should actually do, the output I care about is not a summary. It is a question. ## Science moves on questions, not on facts A field advances when someone notices that two results do not fit together, or that an assumption everyone shares has never been tested, or that a method from another field would answer something this field has given up on. Those are questions, and they hide in the gaps between papers, not inside any single one. A tool that reads one paper at a time will never find them. A tool that reads a thousand and produces a thousand summaries has not helped either. The unit of value is the gap. ## What the instrument should surface Contradictions: two credible sources that report incompatible findings, with the conditions under which each was measured, so the researcher can decide whether the difference is real or methodological. Untested assumptions: claims that are cited constantly and were established once, decades ago, in a population that does not match the current one. Unreplicated results: findings with one source and many citations. Method transfers: an approach that worked for a neighboring problem and has not been applied to this one. Each of these is a structured object with evidence attached, not a paragraph of prose. The output is a list of things worth investigating, ranked by how much evidence points at the gap and how little has been done about it. The researcher reads the list and picks. The instrument does not pick. ## Every item carries its evidence This is the part that is non-negotiable and the part most tools skip. If the instrument says two papers contradict, it has to show the two passages, the populations, the measurements, and the reasoning that classified them as contradictory. If it says an assumption is untested, it has to show the citation chain that leads back to the original source and stops there. A question without an evidence trail is a guess dressed up as insight, and researchers have learned to ignore those quickly. Building this is harder than building a summarizer, because the evidence structure has to be designed before the model touches anything. It is provenance again, applied to reasoning instead of data. ## The human stays in charge of meaning An instrument can say that two findings appear incompatible under certain conditions. Only a scientist can say whether that matters. The difference between those two statements is the entire ethical boundary of the product. I build to the first and refuse to build to the second. Every question the instrument raises is an input to a person who will be accountable for what they do with it. ## Why this is the right first product I have watched what happens to fields with too much literature and not enough time: people read what their advisor read, cite what everyone cites, and the untested assumption survives another decade. An instrument that turns the literature into a ranked list of honest questions attacks that directly. It does not replace the reading. It tells you where to read. That is the kind of acceleration I trust, because the output can be checked, the evidence is visible, and the decision stays with the person whose name goes on the paper. --- ## Postgres is enough, until the day it is not, and how to know that day 2026-06-12 · infra, data My default recommendation for almost any product is a single Postgres instance, well indexed, with a read replica when the reporting queries start hurting. It handles more than people believe. I have seen it carry marketplaces, clinics, logistics platforms and analytics dashboards that would have been three services and a stream processor in a more fashionable design. The boring choice is correct far more often than it is celebrated. But 'Postgres is enough' is not a religion. There is a day when it stops being enough for a specific workload, and the skill is recognizing that day from the numbers instead of from anxiety. ## Signals that are not the day Slow queries are not the day. Slow queries are missing indexes, missing statistics, an ORM producing N+1 patterns, or a table that has never been vacuumed properly. The fix is reading the query plan, which takes an afternoon, not a migration to a different database, which takes a quarter. A large table is not the day. Postgres handles tables with hundreds of millions of rows without complaint if the queries touch them through indexes and the writes are not fighting over the same page. Partitioning by time or tenant extends that further and is still Postgres. Traffic growth is not the day. A well-tuned instance on modest hardware serves thousands of transactions per second. If you are seeing problems at a few hundred, the cause is almost always connection management, a missing pooler in front of a serverless deployment, or lock contention from a hot row, not the engine. ## Signals that are the day The first real signal is write throughput on a single hot path exceeding what one primary can absorb, after connection pooling, after batching, after removing unnecessary indexes on that table. Postgres has one primary. If the sustained write rate on one table is approaching what a single machine's disk and WAL can sustain, and you have already scaled the hardware, that workload needs to be split or moved. The second is an access pattern the relational model fights. Time series appended at high rate and queried only by recent window. Deeply nested documents that are always read and written whole. Graph traversals of variable depth. Full-text search with ranking and faceting at scale. Postgres can do each of these, and there are extensions for most, but when one of them becomes the dominant workload, a specialized store will be cheaper to run and easier to reason about. The third is operational: the instance is large enough that a restore takes longer than the business can be down, or a major version upgrade requires a maintenance window nobody can approve. At that point the question is not 'which database' but 'how do we split state so no single piece is that large'. ## How to know the day in advance - Track transactions per second and WAL bytes per second on the primary, and note the trend, not the value. - Track the ratio of time spent waiting on locks to time spent executing. - Track the top five queries by total time weekly, and notice when the same one keeps winning after optimization. - Track how long a full restore takes, from the last rehearsal, against the recovery time the business expects. - Track connection count against the limit, especially on serverless deployments where every instance opens its own. When one of these crosses a line you defined in advance, you have a decision to make with data. When none of them has, the answer to 'should we move off Postgres' is no, and the next task is an index. ## What moving actually means The day rarely means replacing Postgres. It means taking one workload out of it. The events table goes to a time-series store or a log. The search index goes to a search engine, fed by change data capture. The session cache goes to a key-value store. The core transactional data, the part that has to be correct, stays exactly where it was. Teams that understand this move one thing at a time, keep the boring core, and never need the big migration. Teams that do not either leave too early, into complexity they cannot operate, or too late, in a crisis nobody wanted. The numbers above are how you avoid both. --- ## Designing for deletion 2026-06-10 · architecture, data Nobody designs the delete. Teams spend weeks on the create flow, the validation, the happy path, and then someone adds a DELETE endpoint in an afternoon that runs one statement and returns 204. Six months later that endpoint is the source of an orphaned invoice, a broken foreign key, an analytics table that still counts the customer, and a privacy request nobody can fulfill. I have come to believe deletion is the hardest operation in most systems, and that designing it first is the fastest way to find out what your data model actually is. ## Deletion is a question about ownership Deleting a record forces you to answer a question the create path lets you avoid: what else does this thing own? A customer has orders. Orders have payments. Payments have refunds. Refunds have ledger entries. Does deleting the customer delete the ledger? Obviously not, accountants exist. So what does it do? There are only a few honest answers. Cascade, where the children go with the parent, right for things that have no meaning alone, like line items on an order. Restrict, where the delete is refused while children exist, right for anything with legal or financial weight. Detach, where the children survive with the reference cleared or pointed at a tombstone, right for things like comments on a deleted account. And anonymize, where the record stays but the personal data is scrubbed, which is what most privacy laws actually ask for. Every relationship in the model needs one of those four answers, written down. If you cannot answer for a relationship, you do not understand it yet, and that is exactly the kind of gap that produces the bugs I find in audits. ## Soft delete is not a decision, it is a postponement The common escape is the deleted_at column. Nothing is ever really deleted; a flag is set and every query gets a filter. I use it, but I am honest about what it is: a way to postpone the ownership question, not to answer it. Soft delete has real costs. Every query must remember the filter, and one forgotten filter leaks deleted data into a report or a screen. Unique constraints break, because the deleted email still occupies its row. Foreign keys still point at 'deleted' rows, so the children are in a state nobody designed. And the data is still there, which means the privacy request is not fulfilled, the backup still has it, and the disk still pays for it. When I use soft delete, I pair it with three things: a database view or a default scope that applies the filter so it cannot be forgotten, unique indexes that include the deleted flag, and a scheduled hard delete after a retention period that runs the real ownership rules. Soft delete becomes a grace period, not a final state. ## The delete path is a workflow In any system beyond trivial, deletion is not a statement. It is a workflow with steps that can fail independently: revoke sessions, cancel subscriptions at the payment provider, remove from the search index, purge the CDN cache, notify downstream consumers, scrub backups after their retention, write an audit record that says what was deleted, by whom, and why. So the delete should be modeled like any other long-running operation: a request record with a status, idempotent steps, retries, and a way to see where it is. A 204 that returns before the search index is updated is lying to the user, who will search and find the thing they just deleted. It also means the audit record is not optional. The one thing a deleted record must leave behind is evidence that it existed and was deleted on purpose. Without that, the first support ticket saying 'my data disappeared' has no answer. ## What I check When I audit deletion, I pick one central entity, ask the team to delete a test instance in staging, then look at every table, index, cache and external system that referenced it. Something is always left behind. Usually the search index or an analytics event stream; often a file in object storage; sometimes a webhook a partner already received and cannot un-receive. Then I ask about the privacy request path: a real person asks to be forgotten, what happens? If the answer involves a manual SQL script and a person named in a wiki, the system does not support deletion. It supports someone heroically doing it by hand. Design the delete first. It tells you what owns what, which relationships have weight, which external systems hold copies, and how much of your data model is a story you have not finished telling. Everything you learn makes the create path better too. --- ## Code review is risk review 2026-06-08 · audit, engineering I have read thousands of pull request reviews and most of them review the wrong thing. Naming, formatting, whether a helper should be extracted, whether the tests use the preferred assertion style. All fine, all cheap, all beside the point. The reason we review code before merging it is that merging is a decision to accept a risk, and the review is the only moment where that risk gets a second pair of eyes. If the review does not talk about risk, it is a style check with a merge button. ## The questions that matter When I review, I hold four questions in my head. What can this change break? Not in theory, concretely: which tables, which endpoints, which jobs, which customers. Who finds out first if it breaks, and how: an alert, a support ticket, a quarterly report? How do we undo it: a revert, a feature flag, a data migration backwards, or is this a one-way door? And what is the blast radius if the worst case happens on a Friday night? A change that passes all four is safe regardless of naming. A change that fails one deserves a conversation regardless of how clean it looks. ## Reviewing the diff is not enough The diff shows what changed. Risk lives in what the change touches. A one-line change to a shared utility is higher risk than a two-hundred-line new feature nobody calls yet. A change to a default value affects every caller that did not set it explicitly. A migration that adds a column is safe; one that renames it breaks every query that still uses the old name, including the ones in the reporting system nobody remembers. So I read the diff, and then I read the callers, and then I read the callers of the callers, until I run out of surprise. ## Data changes get a different bar A bug in code is fixed by a deploy. A bug in data is fixed by an engineer, by hand, at night, with a backup open in another window. Any change that writes to data differently than before, alters a schema, backfills, deletes or transforms records gets reviewed as if it were production surgery, because it is. I want to see the query that estimates the affected rows. I want to see the rollback. I want to know it was run against a copy first. A review that approves a data migration in thirty seconds is not a review. ## Reversibility is the cheapest safety The most effective review comment I leave is usually some version of: can we make this reversible? Put it behind a flag. Add the new column before removing the old one. Write the new path alongside the old one and switch readers later. Log before you enforce. Every one of these turns a one-way door into a two-way one, and two-way doors can be walked through with far less scrutiny. Most of the risk in a change is not in the change itself but in how expensive it is to take back. ## What this does to a team Teams that review for risk write different code. They start including the rollback plan in the description because they know it will be asked. They split scary changes into safe steps because a small reversible change gets approved in minutes and a large irreversible one gets a meeting. They stop arguing about naming, because naming can be fixed later and a corrupted table cannot. The review becomes the place where the team decides, together, what it is willing to lose. That is what it should have been all along. --- ## TypeScript generics that help, and the ones that hurt 2026-06-05 · fullstack, typescript Generics are where TypeScript code goes to become either very clear or completely unreadable, and I have reviewed both in the same file. The difference is not skill. It is whether the generic is doing a job that only a generic can do. Here is the test I apply, and it takes five seconds: does the type parameter appear at least twice in the signature? If it relates an input to an output, or two inputs to each other, it is earning its place. If it appears once, it is a costume for unknown, and the function would be clearer without it. ## The ones that help The best generic relates what goes in to what comes out. A function that takes an array of T and returns the first T. A fetcher that takes a schema and returns the schema's inferred type. A repository method that takes an entity name and returns that entity's row type. In each case, the caller writes no annotation and gets a precise type back, because inference carries it through. The second good use is constraining a relationship between arguments. A function that takes an object and a key of that object, typed as K extends keyof T, cannot be called with a key that does not exist. That is a generic doing the work of a runtime check at compile time, and it is the pattern behind every well-typed form library and query builder. The third is a typed container. A Result type with a success branch of T and an error branch of E. A paginated response of T. These are generic because the shape is the same and the content varies, which is exactly what generics were invented for. In all three, the type parameter is inferred at the call site. If callers have to write the angle brackets by hand every time, the generic is probably not helping them. ## The ones that hurt The single-use generic is the most common. A function declared as taking T and returning void, where T is never constrained or related to anything. The author wanted to say 'this accepts anything', and unknown says that honestly. The generic says it with ceremony and implies a relationship that does not exist. The generic that replaces a union is the second. A component that takes a type parameter for its variant, when the variant is one of four known strings. A union of literals is shorter, produces better error messages, and the compiler can check exhaustiveness on it. Generics cannot enumerate. The deeply conditional generic is the third and the worst. Types that recurse through object shapes, distribute over unions, and infer through template literals to compute a result. They are impressive. They are also slow to check, impossible to debug when they fail, and they produce error messages that fill the screen. I have watched a team lose two days to a type that was trying to derive a form's validation shape from its component tree. A hand-written type would have taken ten minutes and been readable forever. The fourth is the generic that leaks. A function returns a T that was never narrowed, so the caller receives a type that is technically correct and practically useless, and reaches for an assertion to make it usable. That assertion is where the runtime bug will be. ## The rules I apply in review A type parameter must be used at least twice, or it becomes unknown. A type parameter must be inferable from the arguments, or the function gets an overload instead. A conditional type more than two levels deep gets a name, a comment with an example of its input and output, and a test using an expect-type helper. Prefer `satisfies` and const assertions to generics when the goal is to check a literal object against a shape without losing its precision. Prefer a discriminated union to a generic when the set of options is known. Use the NoInfer utility when a parameter should be checked against an already-inferred type rather than widening it. And when in doubt, write the concrete version first. Two concrete functions that share a shape will tell you what the generic should be. One abstract function written before the concrete cases exist will tell you what its author hoped the shape would be, and hope is not a type. --- ## Designing APIs people can misuse safely 2026-06-02 · system-design, engineering The client of your API is not malicious. It's worse: it's ordinary. It retries when the network blinks. It sends the same request twice because a user double-clicked. It holds a stale copy of a record for an hour and then writes it back. It calls step three before step two because someone refactored the mobile app. None of that is an attack, and all of it will happen this week. An API is well designed when the ordinary misuse is harmless. That's the bar, and it's mostly a list of decisions. ## Assume every request will arrive twice Any operation that creates or charges something should accept an idempotency key: a client-generated identifier sent in a header. The server stores the key with the result for a window, 24 hours is a common choice, and returns the stored result if it sees the key again. The second click, the retried POST, the replayed webhook all collapse into one effect. Reject a reused key with a different payload rather than silently serving the old response; that is a client bug you want surfaced. For updates, use conditional writes. Return a version or ETag with every read and require it on write: If-Match on HTTP, or a version column in the body. A stale write then fails with 409 or 412 instead of overwriting a newer change. Last-writer-wins is a choice; make it consciously, not by omission. ## Refuse ambiguity, require intent If a field can be interpreted two ways, reject it. A date without a timezone, an amount without a currency, a boolean sent as the string 'false', a null that might mean 'clear it' or 'leave it alone'. Validation errors are cheap. Data corruption from a guessed interpretation is not. Make unknown fields an error in strict endpoints, because a typo in a field name that's silently ignored is the hardest bug to notice. Dangerous operations should require intent. Deleting an account, wiping a dataset, issuing a bulk refund: make them two steps. The first call returns a confirmation token and a summary of what will happen, the second call carries the token. Or require the client to echo the resource's name. A retry of the first call is harmless; the second cannot be reached by accident. Soft delete with a recovery window covers the remaining cases. ## Version explicitly, page reliably Put the version somewhere the client has to choose it: a path segment or a required header. An unversioned API is a version 1 you can never change. When you break something, ship a new version and keep the old one for a stated period, with a deprecation header on every response. Pagination is where misuse hides. Offset pagination over a live collection skips or duplicates items whenever something is inserted or deleted between pages. Use keyset pagination: an opaque cursor encoding the last seen sort key, with a stable sort that includes a unique tiebreaker. The client cannot compute a cursor, cannot skip, and gets each item exactly once even while the table changes. ## Errors that say what to do, limits that say when An error response should say what was wrong, which field, and what to do about it, in a machine-readable shape: a stable code, a human message, the field path. A '400 Bad Request' with no body wastes an hour of someone's day. A 429 should carry Retry-After. A 503 should carry it too. And 5xx should be reserved for things the client cannot fix, so a retry policy can be simple: retry 5xx and 429 with backoff, never retry 4xx. Rate limits belong on every public endpoint, per client and per resource where it matters, with the limits documented and returned in headers. They protect you from bugs as much as from abuse: a client stuck in a retry loop looks identical to an attacker. ## Safe by default Every default should be the one that does the least damage. Lists return a small page, not everything. Omitted filters are inclusive, never destructive. Timeouts are set. Optional booleans default to the conservative behaviour. A new client that sends the minimum should get a safe, boring result. If misusing the API takes effort, you have designed it well. --- ## What AI cannot audit 2026-05-30 · ai, audit I have spent a large part of my career auditing systems, and I use models in that work every day. They read faster than I do, they never get tired on the four hundredth file, and they are very good at pattern matching against known classes of bugs. So I want to be precise about the limit, because it is not where people assume. The model is not weak at finding bugs. It is weak at finding the failures that look like correct behavior, and those are the ones that bring systems down. ## What it does well Hand a model a codebase and it will find the unhandled promise, the SQL built by string concatenation, the missing index, the retry without a limit, the secret in the config. This is real and it is valuable, and I run it early in every audit to clear the known classes so I can spend my time elsewhere. It also summarizes: what does this service do, where does this data flow, which modules talk to which. That map used to take me a day to build. Now it takes an hour to verify. ## What it cannot see The failures I get called for are almost never in the code alone. They are in the gap between what the code does and what the business assumed. A reconciliation job that runs correctly on a schedule that no longer matches when the upstream data arrives. A permission check that is correct for the org structure of two years ago. A cache that is invalidated properly on every write path except the one added by a different team in a different repository. A retry that is idempotent for the API it was written for and not for the one it now calls. To find these you need the intent, the history, and the context outside the repository, and the model has none of them unless a person brings them. ## Plausibility is the enemy The model's strength is generating the plausible reading of the code. Auditing is the search for the implausible truth. When I ask a model whether a function is safe, it explains why the function is safe, fluently, because that is the reading the code supports on its surface. The bug is in the assumption the code makes about its caller, and the model has not seen the caller, and even when it has, the caller's own assumption is about a deploy order documented in a chat message from a person who has since left. A convincing explanation of why the code works is the most dangerous output an audit tool can produce, because it ends the search. ## How I actually use it - As a first pass to clear the known bug classes, so the human time goes to the unknown ones. - As a map builder, and then I verify the map against the running system, not the code. - As a hostile reader: I ask it to argue that a component is broken, not to tell me if it is, and I read the argument for the parts it could not support. - As a documentation check: where its summary of what the code does differs from what the team says it does, there is a finding. ## The accountable reader An audit is a claim signed by a person: I looked, here is what I found, here is what I did not find. A model cannot sign that. Not because it is not smart enough but because the claim includes judgment about what matters in this business, for these users, at this moment, and responsibility for being wrong. That is the same reason ELUCENIA, the global medical and scientific network I am building for discovery, treats every model output as a hypothesis that needs evidence and a human who is accountable for the conclusion. The tool widens what one auditor can cover. It does not replace the auditor. The failures that matter are still found by someone who refuses to believe the plausible story. --- ## Multi-tenancy: the decision you cannot undo 2026-05-27 · system-design, architecture There is one decision in a SaaS product that you make once, early, usually before you have customers, and that shapes every table, every query, every backup and every compliance conversation for the rest of the product's life. It is how you isolate tenants. I have watched teams try to change it later. It is never a migration of one thing. It is a migration of everything. ## The three models Shared schema: every tenant's rows live in the same tables, distinguished by a tenant_id column. Cheapest to operate, easiest to migrate, and one query without a WHERE tenant_id clause leaks data between customers. Schema per tenant: one database, one schema per tenant. Stronger isolation, per-tenant restore becomes feasible, but migrations now run N times and connection pooling gets awkward past a few hundred tenants. Database per tenant: full isolation, trivial per-tenant backup and restore, natural fit for data residency requirements. Operationally it is a fleet: N databases to monitor, upgrade, migrate and pay for. None of these is right. Each one is a different set of pains, and the point is to choose the pains you can live with. ## What actually differs Isolation is the obvious axis, but the practical ones matter more. Noisy neighbours: in a shared schema, one tenant running a heavy report slows every other tenant on the same database, and there is no knob to fix it short of query-level throttling. Per database, the noise stays inside its own walls. Backups and restore: a customer asks you to restore their data to yesterday because an admin deleted everything. With a database per tenant, that is a routine operation. With a shared schema, it is a careful surgical extraction from a full restore, and it takes a day if you are lucky. Compliance and residency: some customers will require their data to sit in a specific country or to be physically separated from other customers. Shared schema cannot offer that without carving out an exception, and the exception becomes a second architecture. Migrations at scale: a schema change in the shared model is one migration. In the per-tenant models it is hundreds, some of which will fail partway, and now your tenants are on different schema versions. ## The tenant_id discipline If you go shared schema, and most products should start there, tenant_id goes on every table. Not just the top-level ones. Every join table, every log, every attachment. And every query filters by it, enforced in the data access layer, not by convention. Then add row level security in the database as a net. Set the tenant on the connection, let the policy filter rows, and a forgotten WHERE clause returns nothing instead of returning everything. I have found the missing WHERE in more audits than I would like to admit. The net catches it. ## The whale Every SaaS eventually lands a tenant that is ten or a hundred times larger than the rest. The whale breaks shared assumptions: its indexes dominate, its reports become the noisy neighbour for everyone, and its compliance team asks for things the others never did. Design for the whale's arrival even if you are not building for it yet. The most practical hybrid is shared schema by default with the ability to move one tenant to its own database, which requires that tenant_id already exists everywhere and that your code never assumes a single connection string. ## Why you cannot undo it Moving from shared to per-tenant means extracting every tenant's rows from every table into new homes, while the product keeps running, and rewriting every piece of code that assumes one database. Moving the other way means merging N schemas and resolving every id collision. Both are multi-month projects with real data-loss risk. So make the choice knowing the trade-off, write it down, and build the escape hatch for the whale on day one. It is the one decision that does not get cheaper with time. --- ## Evidence-first software 2026-05-25 · science, engineering Most software makes claims without evidence. A dashboard says revenue is up. A recommendation engine says you will like this. A risk score says this patient is high priority. Click on any of those and ask why, and the answer is usually silence, or a tooltip nobody can trace to a source. I have come to think this is the deepest design flaw in modern software, and in scientific tools it is disqualifying. ## What evidence-first means The conclusion is the last thing the system produces, not the first. Before a number reaches a screen, the system has to have assembled the inputs that support it, recorded which ones were used, and kept them attached to the output. The user should be able to open any claim and see the trail: this value, from these sources, transformed this way, at this time. If the trail does not exist, the claim does not ship. That sounds expensive. It is cheaper than the alternative, which is that nobody trusts the system and everyone rebuilds the calculation in a spreadsheet to check it. ## The data structure is a claim with attachments In practice I model outputs as claims. A claim has a value, a confidence or status, and a list of evidence items. Each evidence item points to a source, a location in that source, and the transformation that connected it to the claim. This is not exotic. It is a small schema with a foreign key. The discipline is in refusing to create a claim without at least one evidence item, and in making the interface render the evidence one click away from the value. When I audit a system, the first sign of trouble is a table where important numbers are stored without anything pointing back to where they came from. It means somebody computed them once, was confident, and moved on. Six months later nobody can reproduce them and nobody dares delete them. ## Instrument before you decide Evidence-first also applies to the building of the software itself. Before I add a feature that claims to improve something, I want the measurement in place that would show whether it did. Before I optimize a query, I want the trace that shows it is slow. Before I accept that a model output is useful, I want the evaluation set and the baseline. It is the pre-registration habit from science applied to engineering: say what you expect, then look. Teams resist this because it slows the first week. It speeds up every week after, because arguments become lookups. ## Where it matters most In a scientific instrument, evidence-first is not a nice property. It is the difference between a tool and a rumor generator. A system that surfaces a hypothesis about cardiovascular disease must be able to show which papers, which data, which reasoning led there, and it must be honest about what it does not have. A hypothesis without an evidence trail is not accelerating science. It is adding noise to a field that already has too much of it. This is one of the principles I hold ELUCENIA to: every hypothesis needs evidence. Not as a slogan, as a constraint. If the evidence is empty, the claim does not exist. ## The habit transfers Once you build one evidence-first system, you cannot go back. You start looking at every dashboard and asking where the number came from, and you notice how often nobody knows. That discomfort is useful. It is the same discomfort a good scientist feels reading a claim without a citation, and it is exactly the instinct that software has spent two decades training out of its users. --- ## Zero-downtime migrations in practice 2026-05-22 · infra, data Every team says they do zero-downtime migrations. What most of them mean is that their migration tool runs before the deploy and usually nothing breaks. The difference between usually and always is a set of rules about locks, ordering and backfills that are not complicated but must be applied every single time. Here is how I actually do it on Postgres, which is where most of my work lives. ## Rule one: know which statements lock Adding a nullable column with no default is instant. Adding a column with a default was a full table rewrite for many years; on current versions it is instant for constant defaults, but not for volatile ones, and the team must know which version they run. Adding an index without the concurrent option locks writes on the table for the duration of the build. On a large table that is minutes of downtime disguised as a migration. Adding a foreign key validates every existing row under a lock. Adding a not-null constraint scans the whole table. Changing a column's type rewrites it. Renaming a column is instant but breaks every running instance of the old code, which is the subject of rule three. The practical habit: before any migration on a table above a few million rows, read the documentation for that exact statement on that exact version, and set a lock timeout of a few seconds so that a migration that would block waits briefly and fails instead of stalling production behind it. ## Rule two: build indexes and constraints in two steps Indexes are created concurrently. It is slower, it can fail and leave an invalid index that must be dropped and retried, and it does not block writes. Those are the correct trade-offs for production. The migration tool must run it outside a transaction, which most tools support with a flag, and the migration author must know to set that flag. Constraints are added as not valid first, which is instant and applies only to new rows, then validated in a separate statement that scans the table without blocking writes. Not-null on an existing column is done by adding a check constraint as not valid, validating it, and then, on recent versions, setting not null using that constraint as proof. Each of these is two migrations, not one, and the second can run hours later. ## Rule three: expand, migrate, contract Every change that the old code cannot understand is split into phases across separate deploys. Expand: add the new column, table or field, nullable, unused. Deploy code that writes to both old and new. Backfill the new from the old in batches, a few thousand rows at a time, with a pause between batches so replication and vacuum keep up. Deploy code that reads from new and still writes to both. Deploy code that writes only to new. Contract: drop the old column, in its own migration, after a release cycle has passed and rollback to the previous code is no longer plausible. It is slow. A rename that would take one line takes four releases. The alternative is a rename that works in staging, where no old instances are running, and fails in production during the thirty seconds when both versions coexist. ## Rule four: backfills are jobs, not migrations A backfill that touches millions of rows does not belong in a migration file. It belongs in a job that is idempotent, resumable, batched, rate-limited, and observable. It records its progress so it can be stopped and restarted. It is run manually the first time, watched, and only then scheduled. When it finishes, a separate check confirms that the old and new agree before any code starts trusting the new. ## The checklist before every migration - Which statements in this migration take a lock, and for how long on the production table size? - Is every index concurrent, and does the tool run it outside a transaction? - Can the currently deployed code run against the schema after this migration? - Can the previous release run against it, if we roll back? - Is there a backfill, and is it a job with progress tracking rather than a statement? Five questions, asked in review, every time. That is the whole practice. It is not clever. It just refuses to skip the boring part, and the boring part is where the downtime lives. --- ## The full-stack engineer is a systems engineer 2026-05-20 · fullstack, career The term full stack has been diluted into a job title that means 'knows React and can write an endpoint'. I want to argue for the original meaning, because I think it describes a different and more valuable person. A full-stack engineer is someone who can be handed a broken product and find the failure wherever it is: in a CSS layout shift, in a misconfigured cache header, in an N+1 query, in a DNS TTL, in a race condition between two workers. Not because they are an expert in each layer, but because they refuse to stop at the layer boundary and say 'not my problem'. That is a systems engineer. The stack is just where the system happens to live. ## The boundary is where the bug hides In six hundred plus projects, I have almost never found a serious failure that lived cleanly inside one layer. It lives between them. The frontend assumes the API returns a sorted list and the database stopped guaranteeing it after an index change. The API assumes the queue is at-least-once and the consumer assumes exactly-once. The cache thinks the session is valid and the auth server revoked it two minutes ago. A specialist in each layer can look at their piece, confirm it is correct, and be right. The bug persists. It takes someone willing to hold the whole request path in their head to see that the pieces are individually correct and jointly wrong. That is the actual skill. Not breadth of syntax. Breadth of responsibility. ## What the job really requires I can name the competencies, because I have watched engineers grow into them. You need to read a network waterfall and know which request is the bottleneck. You need to read a query plan and know why the index is not used. You need to know what a load balancer does to sticky sessions, what a CDN does to Set-Cookie, and what happens to your Server Action when the deploy finishes mid-request. You need to know enough infrastructure to ask the right question of the person who knows more. That includes the boring parts: environment variables, TLS termination, log shipping, what a health check should actually check. And you need product sense, because a system that is correct and unusable is still a failure, and the full-stack engineer is usually the one closest to both the user and the database at the same time. ## Why teams should want this person The economic case is simple. Handoffs are where time goes. A feature that needs a frontend engineer, a backend engineer and an infra engineer to coordinate takes three calendars to schedule and three mental models to align. One person who can own it end to end ships it in a fraction of the elapsed time, and there is nobody to blame across a boundary. The quality case is stronger. When one person owns the request path, the failure modes get designed instead of discovered. They know the form will be resubmitted, because they wrote the handler. They know the query will be slow at scale, because they wrote the schema. The feedback loop is inside one head. This is not an argument against specialists. A serious product needs a database expert, a security expert, a designer. It is an argument that the person integrating their work must understand all of it well enough to know when it does not fit. ## How to become one You do not become a systems engineer by learning another framework. You become one by following a bug past the point where it stopped being your job. Next time a request is slow, do not file a ticket for the backend team. Open the query plan. Next time the deploy breaks, do not wait for infra. Read the pipeline. I am self-taught, and the way I learned every layer was by being the only person available when it failed. That is not a plan I recommend, but the principle holds: the stack is not a list of technologies. It is the set of places your product can break, and your job is all of them. --- ## What a doctor taught me about systems 2026-05-19 · life, system-design I met my father once. I was a teenager, he was in his seventies, and we had about two hours in the clinic in Campo Limpo where he had worked for more than forty years. He asked what I wanted to be. I said lawyer. He died a few months later. I did not become a lawyer, and I did not become a doctor either, but I have spent most of my adult life doing something close to what he did in that room: looking at a thing that is not working, finding out why, and fixing the cause instead of the symptom. I want to write down what a doctor's way of working taught me about systems, because it holds up better than most of the engineering advice I have read. ## The history comes before the exam A good doctor does not start with the machine. He starts with the story. When did it begin, what changed, what did you do about it, what happened then. The diagnostic tools come after, and they confirm or reject a hypothesis that already exists. When I audit a system I do the same. Before I read a line of code, I ask for the history. When did the slowdowns start. What was deployed that week. Who touched the database. What did the team try, and what happened after each attempt. Most of the time the story narrows the search to two or three places. The logs then confirm which one. Engineers who skip the history spend days in dashboards looking for a signal the team could have told them in ten minutes. ## Symptoms are not the disease Fever is not an illness. It is the body's response to one. Treating the fever without asking what caused it is how you feel better on Tuesday and worse on Friday. In systems, the fever is the retry storm, the growing queue, the CPU pinned at 100%. Teams add capacity, the number goes down, and everyone relaxes. Three weeks later it is back, bigger. The cause was never the capacity. It was a query without an index, a lock held too long, a client that retried without backoff. When I find a metric that "fixed itself" after scaling, I treat that as evidence that the disease is still there. ## The cheapest test first My father worked in a neighbourhood where many patients could not pay, and he treated them anyway. That constraint shaped his method. You do not order the expensive exam when a careful question and a stethoscope will tell you the same thing. You order it when the cheap tests have narrowed the field and you need certainty. In systems this translates directly. Before the distributed tracing project, before the observability platform, before the consultant, read the slow query log. Count the connections. Check whether the timeout is actually configured or just documented. In 600+ projects I have seen more outages explained by a config file than by anything a tracing tool would have found. The expensive instrument is worth it, but it comes after the cheap questions, not instead of them. ## Do no harm includes the fix A surgeon knows that every intervention carries its own risk. The question is never only "will this help" but "what does this cost the patient, and is it worth it". The equivalent for us is the fix that becomes the next incident. The hotfix at 2 a.m. that adds a cache without an invalidation strategy. The migration that runs on the whole table because it "should be quick". The rewrite that was supposed to remove the bug and instead removed the behaviour three other teams depended on. I now ask, before every change on a live system, what is the blast radius if this is wrong, and can I undo it in less time than it took to apply. If the answer to the second question is no, I find a smaller change. ## What I took from those two hours I do not remember most of what we talked about. I remember that he listened before he spoke, that he asked short questions, and that the room was full of people who trusted him because he had been there for forty years and had not left when it was hard to stay. That is the standard I try to hold when someone hands me a system nobody understands anymore. Listen to the history. Look for the cause, not the fever. Try the cheap test before the expensive one. Make sure the cure is not worse than the disease. And stay long enough to see whether the patient actually got better, because the number going down on a dashboard is not the same thing as the system being well. --- ## The human in the loop is a design decision 2026-05-15 · ai, science "There is a human in the loop" is the sentence I hear most often when a team presents an AI feature with real consequences. It is meant to end the discussion. For me it starts one. A human in the loop is a design decision with a dozen sub-decisions, and most of the time nobody has made them. The human exists on the architecture diagram as a box labeled "review". In production, that box is a person clicking approve forty times an hour without reading. ## Not a checkbox Putting a person between the model and the consequence does not make the system safe. It makes the system safe only if the person can actually catch what the model gets wrong. That requires three things that do not happen by accident: the person sees the information needed to judge, the person has time and incentive to judge, and the person's judgment changes what happens. Remove any one of them and the human is a rubber stamp with a salary. I have audited approval flows where the reviewer had a quota, a ten-second average and no view of the source. The loop had a human. The system had no review. ## Three questions When I design a human step, I ask what the human is for. Is it to catch errors, in which case they need to see the model's evidence and its uncertainty, not just its conclusion? Is it to take responsibility, in which case the interface must make clear what they are signing and record that they did? Is it to teach the system, in which case every correction must flow back into the eval set and the next version? These are different jobs with different interfaces. A single approve button serves none of them well. ## Where the human adds value The human should be placed where their judgment is scarce and the model's is weak, not sprinkled everywhere as insurance. Reviewing a thousand routine extractions is a job for validation code and sampling. Deciding whether an unusual case should override the policy is a job for a person. Good designs route by confidence and consequence: the model handles the confident, low-stakes majority alone, code checks the mechanical properties, and the person receives the small set where the model is unsure or the outcome is irreversible, with everything they need to decide. That is a queue with a priority, and it can be measured like one. ## Interfaces for accountability The part most teams skip is the record. Who saw what, when, and what did they decide. Not for blame, for learning. When an approved output turns out to be wrong, the question is whether the reviewer could have known, and the answer depends on what the interface showed them. If the interface hid the model's low confidence, the failure belongs to the design, not the person. So I log the full context of every human decision, and I periodically review the decisions themselves, because an approval rate near one hundred percent is not a sign of a good model. It is a sign that the loop has stopped looping. ## Science This is not abstract for me. The reason I am building ELUCENIA, a global medical and scientific network for discovery, is that I believe machines can shorten the path from data to hypothesis. But a hypothesis is not a result. Every hypothesis needs evidence, every advance needs validation by the scientific method, and every decision needs a human who is accountable for it. Those are not limits on what the tool can do. They are what makes the tool trustworthy enough to be used at all. The human in the loop is the design. The model is the assistant. --- ## The shared library trap 2026-05-13 · architecture Somewhere in almost every multi-service codebase I audit there is a package called common, core, shared or utils. It started as a good idea: three services needed the same date formatting, the same auth middleware, the same error type. Copying it three times felt wrong. So someone extracted it, published it, and everybody depended on it. Two years later, that package is the reason nothing can be deployed on a Friday. It has forty dependencies. It contains the database models. Bumping it in one service means bumping it in twelve, and one of the twelve is owned by a team that no longer exists. The library created to reduce coupling became the single tightest coupling in the system. ## How the trap closes The trap has a predictable shape. First, the shared package holds only pure helpers: formatting, validation, small types. That is fine. Then someone adds a client for an internal API, with its retry policy and its config. Then the ORM models go in 'so they are in one place'. Then a feature flag reader that needs a network call at import time. At each step the package gets more useful and less stable. Each addition drags in dependencies and assumptions about the runtime: an environment variable that must exist, a database that must be reachable, a specific framework version. The consumers did not ask for any of that. They wanted the date formatter. The moment I know the trap has closed is when a team pins the shared package to an old version and refuses to upgrade. Now there are two versions in production with different behavior, and the 'single source of truth' is two truths that nobody compares. ## Duplication is cheaper than the wrong abstraction The instinct to extract comes from a rule taught early: do not repeat yourself. It is a good rule inside one codebase with one deploy unit. Across services with independent deploys, it inverts. Every shared line is a line that must change in lockstep, and lockstep is exactly what services were supposed to escape. My threshold: if two services have similar code, leave it. At three, look closely and ask whether it is the same because of a shared concept or by coincidence. Date formatting is coincidence; both could change independently and nothing would break. A money type with currency and rounding rules is a shared concept; if it drifts, invoices disagree. Even for shared concepts, the shared thing should often be a specification, not a package. A JSON schema, an OpenAPI document, a protobuf file. Each service generates or writes its own implementation against the spec. The spec is versioned and stable. The implementations are free. ## What belongs in a shared package There is a shared package that works, and it is boring. It has zero runtime dependencies, or as close as the language allows. It does no I/O. It reads no environment. It has no opinion about which framework you use. It is versioned semantically, and every breaking change ships with a codemod or a clear migration note. Things that pass that test: pure types, value objects like Money or Email, pure validation functions, constants that really are constant, and tiny algorithms like an ID generator. Things that fail it: HTTP clients, database models, auth middleware, logging setup, config loaders and anything that knows the name of another service. For the failed group, the alternatives are better than they look. An HTTP client for an internal API belongs to the API's team, generated from its spec and released on the API's schedule. Auth middleware belongs at the edge, in a gateway, not repeated in every service. Logging setup is ten lines; copy it. Database models belong to exactly one service, the one that owns the table, and everybody else talks to it through an API or an event. ## Getting out If you are already in the trap, do not attempt a big bang. Start by measuring: for each export in the shared package, count the consumers. In my experience a third of the exports have one consumer and can move back into it today. Another third are pure and can stay. The last third are the problem, and they are usually the ones doing I/O. For those, pick the most painful, find its rightful owner, and move it there with a deprecation shim in the shared package that re-exports from the new home. Remove the shim after two release cycles. Repeat. It is slow, which is fine, because the alternative is a package nobody dares to touch. The test I use at the end: can any single service upgrade the shared package, alone, on a Tuesday, without a meeting? If yes, the library is a library. If not, it is a distributed monolith with extra steps. --- ## Details are not small 2026-05-11 · audit, life People describe me as detail-obsessed and they usually mean it as a mild criticism. Too slow, too thorough, unwilling to let the nullable column go. I understand the frustration. I also think the word 'detail' is doing a lot of dishonest work, because it implies the thing is small. It is not small. It is a decision, made at a small scale, whose consequences are not. ## A default value is a policy Take the most boring example I know: a default value on a field. The column is created with a default of zero, or the API parameter defaults to false, or the form pre-selects the first option. That default will be applied to every record where nobody thought about it, which is most records. It is a policy, silently enacted, affecting more rows than any explicit decision the team ever made. I have seen a default timezone set to the developer's own city corrupt years of scheduling data across a whole country. Nobody called it a detail after that. ## Nullable means nobody decided Another one: a column that allows null. Every nullable column is a question the team declined to answer. Can a user exist without an email? Can an order exist without a total? Sometimes the honest answer is yes, and the null means something specific. More often the null is there because the migration was easier that way, and now every reader of that column has to guess what absence means. The bug arrives months later, in a report that quietly excluded every row where the value was missing, and the number everyone made decisions on was wrong by a third. ## Why I refuse to let them go When I audit a system, the findings people push back on hardest are always the ones that look small. A missing index. An inconsistent rounding. A timestamp without a timezone. I hold the line on these not because I enjoy it, but because 600+ projects taught me that the catastrophic failure is almost never a big, visible mistake. It is a small decision that nobody reviewed, compounded across every record and every day, until the day it finally mattered. The big mistakes get caught. The details ship. ## What my father's patients remember I met my father once, for about two hours, as a teenager. Everything else I know about Pedro Moretti Guedes, a doctor in São Paulo for more than forty years, comes from his patients. What they remember is not the big things. It is that he remembered their mother's name. That he noticed when someone had walked instead of taking the bus, and asked why. That he did not charge those who could not pay, and did it in a way that let them keep their dignity. Details. Every one of them a small decision that told a person they were seen. I did not learn that from him directly. I learned it from the people who carried it for decades after. And it shaped how I think about the work: the detail is not the small part of the job. It is the part where you decide whether you actually care. ## Attention is a form of respect In software, attention to detail is respect for the person who will use the system, the person who will maintain it, and the person whose data it holds. A well-chosen default respects the user who will never change it. A non-nullable column respects the analyst who will query it in two years. A timezone-aware timestamp respects the patient in another state whose appointment depends on it. None of these will be noticed when they are right. All of them will be felt when they are wrong. So I keep being slow, and thorough, and unwilling to let the nullable column go. It is not obsession. It is the job, taken seriously, at the scale where it is actually decided. --- ## The first three alerts every system needs 2026-05-08 · infra, reliability Most systems I audit have one of two alerting setups. Either there are no alerts, and the team learns about outages from customers, or there are two hundred alerts, and the team has learned to ignore all of them. Both are the same failure. The alert that nobody trusts is the alert that does not exist. So when a small team asks me where to start, I do not say 'define your SLOs'. That comes later. I say: set up three alerts, make them precise, make them page someone, and refuse to add a fourth until those three have proven themselves for a month. ## Alert one: the front door is failing The first alert is the error rate at the edge. Not per service, not per endpoint. The proportion of user-facing requests that return a 5xx, measured over a short window, compared against a threshold that reflects real damage. The details matter. Use a ratio, not a count, so that traffic growth does not cause false alarms and a quiet night does not hide a real one. Use a window of a few minutes, long enough to smooth a single bad deploy instance, short enough that a customer has not yet written an angry email. Set the threshold by looking at a normal week: if your baseline is 0.1 percent, alert at 1 percent sustained for five minutes, and adjust after a month of data. Exclude the errors that are the client's fault, which means 4xx stays out, and exclude health check endpoints, which would otherwise dominate the count. This one alert catches the majority of outages: bad deploys, database down, dependency down, certificate expired, out of memory. It is not diagnostic. It just says 'something is wrong and users feel it', which is exactly what you want to hear first. ## Alert two: the front door is slow The second alert is latency at the edge, and specifically the tail. The p99 of user-facing request duration over a similar window, above a threshold that represents 'the product feels broken'. Averages hide everything. A p50 of 120 milliseconds is compatible with one in a hundred users waiting eight seconds, and those users are the ones with the biggest accounts and the most data. A p99 threshold set at two or three times your normal p99 catches the slow database, the lock contention, the dependency that is not dead but is dying, the memory pressure that is not yet an out-of-memory kill. Slowness is the early warning for most of the failures that alert one will eventually catch. If you fire on latency, you get to act ten minutes before you would otherwise be paged for errors. ## Alert three: the thing that should be happening is not The third alert is different in kind. It is the absence alert. Every system has a heartbeat: orders being created, jobs completing, messages being consumed, backups finishing. The third alert fires when the heartbeat stops. Error rate and latency both assume traffic is arriving. But if the queue consumer silently crashed, no errors are produced and no latency is measured. The work just stops. If the nightly backup did not run, nothing is red. If the cron job that sends reminders has not executed in two days, no dashboard shows it. Pick the one business process whose silence would hurt most and alert on 'zero completions in the last N minutes, during hours when N minutes of silence is abnormal'. This is the alert that catches the failures the other two cannot see, and in my experience it is the one that finds the incidents nobody else finds. ## Rules for keeping them honest - Every alert pages a human who can act. If nobody can act, it is a dashboard, not an alert. - Every alert has a one-line runbook: what to check first, what to do second. - Every false positive is fixed within the week, by changing the threshold or the query, never by muting. - No new alert is added until it is tied to a concrete failure you have seen or credibly expect. Three alerts, tuned carefully, are worth more than three hundred added carelessly. Start there. Add the fourth when the first three have earned your trust, and not before. --- ## How I read a system I have never seen in one hour 2026-05-06 · system-design, audit I have been dropped into hundreds of codebases with no context and a deadline. Over time I stopped reading code first. Code tells you what the system does on a good day. What I need to know is what it does on a bad one, and where the bad day will start. This is the hour I run, in this order. ## Minutes 0 to 15: the shape First, entry points. Where does traffic come in? HTTP routes, queue consumers, cron jobs, webhooks, admin tools. I list all of them, not just the ones in the README. Cron jobs and webhook handlers are where the surprises usually live, because nobody reviews them. Second, data stores. Every database, cache, object bucket, search index and third-party system that holds state. For each one I want to know one thing: who writes to it. If two services write to the same table, I have found my first architectural risk before opening a single file. Third, the biggest table. Ask for row counts. The biggest table is usually the one with the slowest query, the missing index, and the migration nobody wants to run. It is also where the growth story of the company is written. ## Minutes 15 to 35: the write paths Now I follow writes, not reads. Reads are forgiving. Writes are where the damage is. I trace the two or three most important write paths from the entry point to the disk: create order, charge a card, change a permission, send a message. Along the way I look for the longest transaction. A transaction that opens, calls an external API, and then commits is a lock held for as long as the API takes. That is a deadlock waiting for traffic. Then I ask where money or irreversible actions happen. Payments, refunds, emails, SMS, deletions, anything that leaves the system and cannot be recalled. For each one: is it idempotent? What happens if this runs twice? What happens if it runs once and the record of it is lost? Those two questions find more real bugs than any static analysis tool I have used. ## Minutes 35 to 50: the bad day For every dependency on my list, I ask one question: what happens when this is down? Not degraded, down. The answer should be specific. "Checkout fails but browsing works" is an answer. "It should be fine" is not. Then I open the logs and alerts. Not to read them, but to see what exists. Is there an alert for error rate? For queue depth? For the payment provider returning errors? If the only alert is "server is unreachable", the team learns about incidents from customers. Then the on-call runbook. If there is no runbook, the runbook is one person's memory, and I want to know who that person is and whether they are on vacation. ## Minutes 50 to 60: the questions I end by asking the team a short list, and I listen more to the pauses than the answers. - What was the last incident, and what changed after it? - Which part of the system are you afraid to touch? - What runs at 3 a.m. and who knows what it does? - If the database was restored from last night's backup, what would be lost, and would you know? - Which customer would you call first if it went down? The pause before "which part are you afraid to touch" is where the real audit begins. Everything before was me learning the map. That question is where the team hands me the territory. One hour is not enough to fix anything. It is enough to know where to look for the next forty. --- ## Error boundaries, end to end 2026-05-01 · fullstack, reliability React gave us error boundaries and most teams stopped there. A component throws, a fallback renders, the rest of the page survives. That is one layer of a system that has six, and in audits I find that the other five have no boundary at all. An unhandled rejection in a queue worker takes down the process. A failed webhook retries forever. A Server Action throws and the user sees a generic message with no idea what to do. An error boundary is a decision about where a failure stops propagating and who gets told. That decision has to be made at every layer, and it has to be consistent, or the system fails in the layer you forgot. ## Expected and unexpected are different things The first decision is the classification. An expired card, a validation failure, a record that no longer exists: those are expected. They are part of the domain, they happen every day, and they should be values, not exceptions. A function that can fail in an expected way returns a result type with an explicit error branch, and the caller handles it in normal control flow. Unexpected errors are the database being unreachable, a null where the type said there could not be one, a third-party API returning HTML instead of JSON. Those are thrown, because nobody at the call site can do anything useful about them, and they need to travel up to a boundary that knows how to fail safely. Mixing the two is the source of most bad error handling I read. Throwing for a declined card means every caller needs a try block. Returning a value for a database outage means the outage gets swallowed in a branch that logs and continues. ## The React layer In the App Router, an error file next to a route segment becomes a boundary for that segment. It is a Client Component that receives the error and a reset function. The layout above it keeps working, which is exactly what you want: a broken sidebar widget should not kill the checkout. In production, the framework strips the message from errors thrown in Server Components and sends a digest instead, to avoid leaking internals. That means the boundary cannot show the user what went wrong, and it should not try. Its job is to show a way forward: retry, go back, contact support with the digest attached. The actual message lives in your logs, keyed by that digest. Server Actions are different. They are called from the client, and if they throw, the client gets the same sanitized error. So expected failures in an action are returned, never thrown, as a structured result the form can render field by field. Unexpected ones are thrown, logged on the server with the digest, and caught by the nearest boundary. ## The process layer Below React, there is a Node.js process, and it has its own boundaries. An unhandled promise rejection should crash it, on purpose, because a process in an unknown state serving requests is worse than a restart. The orchestrator restarts it. The health check fails during the restart, and the load balancer routes around it. That only works if the health check is real and the orchestrator is configured, and I verify both before I trust a crash-on-error policy. Queue workers need a boundary per message. A message that fails is retried with backoff a fixed number of times, then moved to a dead-letter queue with the error attached. Without that, one poison message blocks the queue forever, and I have found that exact situation in production more than once. ## The reporting layer Every boundary reports before it recovers. The React boundary sends the digest and the route. The action logs the input shape and the stack. The worker logs the message id and the attempt count. All of it goes to one place, with a correlation id that ties the click to the query. If a boundary swallows an error without reporting, it has not handled the error. It has hidden it. The difference is whether the on-call engineer learns about the failure from a dashboard or from a customer. Draw the boundaries on the whiteboard, one per layer, with an arrow to where each one reports. If a layer has no arrow, that is where the next incident starts. --- ## What a green build hides 2026-04-30 · audit, infra A green build is the most reassuring lie in software. It means the tests that exist passed, in the environment they ran in, against the mocks they were given, on the code that was committed. Every clause in that sentence is a place for reality to differ. I have audited systems with a hundred percent green history and a production that fell over weekly. The build was not wrong. It was answering a narrower question than anyone realized. ## Tests that do not test Open the test file for the most important module and read the assertions. Not the count, the content. I routinely find tests that assert the function returned something, tests that call the function and check nothing, tests that mock the thing under test, and tests that were marked skipped in a hurry and never unmarked. Coverage tools count all of these as coverage. A suite with a thousand tests and no meaningful assertions is a slow way to compile. ## Mocks at exactly the boundaries that fail Tests mock the database, the payment provider, the queue, the clock, the file system. That is reasonable for speed. It also means the test suite never exercises the boundaries, and the boundaries are where the bugs are. The mock returns the shape the developer expected. The real dependency returns a null in a field the mock never had, a timeout the mock never produced, a duplicate the mock never sent. Green means the code works against an idealized world. The question is whether anything, anywhere, runs it against the real one. ## Flaky tests and the retry that hides them Look at the pipeline configuration for a retry-on-failure setting. If the build re-runs failed tests until they pass, the suite has intermittent failures that somebody decided to tolerate rather than investigate. Intermittent failures are almost always real: a race condition, a test that depends on order, a shared fixture that leaks state. The retry turns a signal into silence. When I find one, I turn it off for a day and watch what falls out. It is never nothing. ## The environment is not production The build runs on a clean container with the dependencies from the lockfile and an empty database seeded by fixtures. Production runs on instances that have been up for weeks, against a database with years of data, behind a load balancer, with environment variables that were set by hand in a console and do not appear in any file. Migrations that succeed on an empty database fail on a large one. Queries that are fast on fixtures are slow on real volumes. Feature flags that default one way in test default another in production. The build cannot see any of this, and a green result says nothing about it. The last one is mundane and I have seen it twice in a single year: the green build was not of the code that got deployed. A build cache served a stale layer. An artifact from a previous run was promoted by mistake. The deploy pulled from a branch that was three commits behind. Everything in the dashboard was green. The code in production had never been tested at all. The fix is to make the deployed artifact carry its commit hash and to check it, mechanically, after every deploy. ## What green should mean None of this is an argument against CI. It is an argument for knowing what your CI is actually checking. The exercise I run with teams is simple: write down, in one sentence, what a green build proves. Then write down what it does not. The second list is always longer, and it is the list of things that need a different kind of check: a staging environment with real data volumes, a smoke test after deploy, an integration test that hits the real dependency once a day, a restore drill for the backup, a human reading the assertions. Green is necessary. It was never sufficient. --- ## File uploads, done right the first time 2026-04-28 · fullstack, engineering File upload is the feature every product needs and almost every product gets wrong the first time. I know because I have audited the second time: the server that ran out of memory on a large PDF, the bucket with public read on medical documents, the filename that contained a path traversal, the image that was actually an executable. The correct design has been the same for a decade, and it starts with one rule: the file never touches your application server. ## Upload direct to storage The client asks your server for permission to upload. Your server checks the session, decides whether this user may upload this kind of file of this size, and returns a presigned URL from the object storage provider, valid for a few minutes, bound to a key your server chose and a content length limit. The browser then uploads straight to storage. Your server never sees the bytes. This solves the memory problem, because no Node.js process buffers a gigabyte. It solves the timeout problem, because a Server Action with its default one-megabyte body limit was never going to carry a video anyway. And it solves the scaling problem, because storage providers are built to absorb uploads and your web tier is not. The key is generated by your server: a random identifier, never the user's filename. The original filename is stored as metadata, escaped, for display only. A filename is user input, and I have seen dots and slashes in it do things nobody intended. ## Record first, verify second Before returning the presigned URL, your server writes a database row for the upload in a pending state, with the key, the owner, the declared type and size, and a timestamp. The upload is not real until a second step confirms it. That second step runs after the client reports completion, or after a storage event notification arrives, and it runs asynchronously. It reads the first bytes of the object and checks the magic number against the declared type, because the extension and the content-type header are both claims the client made. It checks the actual size. It runs a malware scan if the product handles documents from strangers. Then it flips the row to ready, or deletes the object and marks the row as rejected. Until the row says ready, nothing in the product references the file. That one rule prevents the whole class of bugs where a half-uploaded or malicious file appears in a list because the client said it was done. ## Serve through a door you control The bucket is private. Always. A file is served either through a short-lived signed download URL that your server issues after an authorization check, or through a route that streams it with the right headers. Public buckets are how private documents end up indexed by a search engine. Every response carries a Content-Disposition header that sets the download filename to something you chose, and a content type that you verified rather than the one the client declared. For user-provided HTML or SVG, the content is served from a separate origin or forced to download, because a file that renders in the browser under your domain is a cross-site scripting vector. Images get processed into your own formats and sizes by a worker, and the product serves the processed versions. The original is kept for reprocessing and is never served directly. ## The details that bite later Orphaned objects are the slow leak. A user requests an upload, closes the tab, and the pending row and the empty key sit there forever. A lifecycle rule on the bucket that deletes objects with no confirmed row after a day keeps the bill honest. Large files use multipart upload, which lets the browser resume after a dropped connection instead of restarting. Above a hundred megabytes, that is the difference between a feature that works on mobile and one that does not. Quotas are enforced at the permission step, before the URL is issued, per user and per tenant. Enforcing them after the upload means paying for storage you are about to reject. None of this is hard. It is a sequence of decisions that have to be made in the right order, and the wrong order is the one that seems simplest on day one: accept the file, save it, deal with the rest later. Later is the audit. --- ## Observability for LLM applications 2026-04-24 · ai, observability The first LLM feature I put in production had the same observability as the rest of the service: request logs, latency histograms, error rates. Everything was green for a week while the feature quietly gave worse and worse answers, because the provider had changed something on its side and none of my metrics could see quality. Traditional observability tells you whether the system is up. For a model, up is the easy part. The question is whether it is right, and how much it cost to be right, and you have to build the instruments for that yourself. ## The trace is the unit For a request that touches a model I record the full trace: the final prompt after all assembly, the retrieved context and its sources, every model call with input and output tokens, the raw output before parsing, the parsed result, the validation outcome, any retry, and the total latency broken into segments. Not a summary. The actual strings. Storage is cheap compared to the hour you spend guessing what the prompt looked like when a user reports a bad answer. Redact what must be redacted, keep the rest, and give it a retention policy. ## Four signals beyond uptime Cost per request, in tokens, sliced by feature and by user tier, because one prompt change can double a bill without touching an error rate. Quality, measured by whatever you can measure automatically: schema validity, retrieval hit rate, a small model judging outputs on a sample, user feedback when you have it. Drift, meaning the distribution of inputs and outputs over time, because a model that starts producing longer outputs or more refusals is telling you something changed. And the fallback rate: how often you served the degraded path instead of the model. When any of these moves, I want an alert, and I want the trace that moved it one click away. ## Replay is the killer feature The single most valuable capability I have built on top of these traces is replay. Take a stored trace, run the same prompt against a new model version or a new prompt template, and diff the outputs. It turns "we think the new prompt is better" into a table of before and after on real traffic. It also makes incident response possible: when a user reports a wrong answer I do not reconstruct it, I open the trace, replay it, and see the failure. Replay requires that the trace captured everything the call depended on, including retrieval results and the system prompt version, so design for it from day one. ## Sampling and privacy You cannot judge every output with a model, and you should not store every prompt forever. I sample for quality scoring, weighted toward the cases that matter: low confidence, retries, long outputs, new users. I hash user identifiers and strip obvious personal data from stored prompts before they leave the request path. And I make it visible to the team that these logs exist and what is in them, because an observability system that quietly stores user conversations is a liability waiting for a subpoena. ## What the dashboard looks like Mine has cost per feature per day, p50 and p95 latency with the model call separated, validation failure rate, fallback rate, the quality sample score, and a stream of the worst-scored traces from the last hour. That last panel is where I look first. A model does not fail loudly. It fails one answer at a time, politely, in complete sentences, and the only way to see it is to go looking. --- ## Architecture decision records that actually get read 2026-04-22 · architecture, engineering Every team I audit has one of two things: no architecture decision records at all, or a folder of forty ADRs that nobody has opened since the day they were merged. The second is only slightly better than the first, because it creates the illusion that decisions are documented. An ADR earns its existence when a new engineer, eighteen months from now, is about to undo a decision and finds the record before they do. That is the only use case that matters. Everything about the format should serve it. ## Why most ADRs die They die because they are written for the wrong reader. The typical ADR is written to justify the decision to the people in the room that week. It uses their vocabulary, assumes their context and skips the alternatives that were 'obviously' wrong. Eighteen months later, none of those people are in the room, the vocabulary has shifted and the obviously wrong alternative is the one the new engineer is about to pick. They also die because they are too long. A three-page document with a 'Context' section that restates the whole product will not be read under pressure. And ADRs are always read under pressure, during an incident or a rushed refactor, never on a calm Tuesday. And they die because they have no expiry. A decision made when the system had a thousand users is presented with the same authority as one made last month. The reader cannot tell whether it still applies. ## The format that survives I keep ADRs to one screen. Title as a full sentence stating the decision, not the topic: 'Orders are stored in Postgres, not in the event store' rather than 'Order storage'. Then five short sections. The decision, in two or three sentences, written so it can be quoted. The forces, meaning the three or four constraints that made this the right call: a number, a deadline, a team size, a compliance rule. The alternatives rejected, one line each, with the specific reason each one lost. Not 'MongoDB was considered' but 'MongoDB was rejected because we need multi-row transactions for refund reconciliation'. The consequences you accept, meaning what gets harder because of this choice. And the conditions for revisiting: 'if order volume exceeds ten thousand per minute' or 'if we add a second region'. That last section is the one almost nobody writes and the one that makes the whole record useful. It turns the ADR from a justification into an instrument. The future reader does not have to guess whether the decision still holds. They check the conditions. ## The discipline around the format The format is half of it. The other half is where the ADRs live and how they are linked. They belong in the repository, next to the code, in plain Markdown, numbered sequentially. Not in a wiki, not in a document tool that requires a login the new engineer does not have yet. Every ADR gets referenced from the code it governs. A comment at the top of the module: 'See ADR-014 for why this does not use the shared cache'. That is the retrieval path. Nobody browses an ADR folder. They arrive at the code first, and the code must point them to the record. Superseded ADRs are never deleted. They get a one-line header, 'Superseded by ADR-031', and stay in place. The history of why a decision changed is often more valuable than the decision itself, because it shows which forces moved. And ADRs are written before the decision is implemented, not after. Writing the alternatives section honestly, while the alternatives are still possible, is what makes the record trustworthy. A post-hoc ADR is a press release. ## What I look for in a review When I audit a system, I read the ADR folder before the code. If the folder is empty, I know every major decision is in someone's head. If it is full but no record has a 'revisit when' section, I know the decisions were written to close an argument, not to inform a future one. If the code never references them, I know they are not part of the working system. The good sign is small and specific: a record from two years ago that someone updated last quarter with a note saying the condition was met and here is the follow-up decision. That is a team that treats architecture as a living thing with a memory. It is rare. It is worth building. --- ## Self-taught is not alone-taught 2026-04-22 · career People call me self-taught, and it is true that I never finished a degree in anything related to software. But "self-taught" gives the wrong picture. It suggests a person alone in a room with a laptop, absorbing everything from documentation. That is not how I learned, and in my experience it is not how anyone learns well. I learned from Samuel, who ran a print shop and taught me CorelDRAW and vinyl cutting when I was about twelve. I learned from Val, who repaired electronics and showed me how to find a fault on a board by thinking about where the current goes. Neither of them was a software engineer. Both of them taught me things I use every week. The self-taught engineer is not someone who learned alone. It is someone who had to find their own teachers, because nobody assigned them. ## What a teacher actually gives you Samuel did not teach me CorelDRAW. The software I could have figured out. What he taught me was that the file you send to the plotter must be exactly right, because vinyl is not free and a wrong cut is money on the floor. He taught me to check twice before committing. Val taught me that you do not replace components at random until the thing works. You measure, you reason about the circuit, you narrow down, and then you replace one thing. That is debugging. I did not know the word yet. What a teacher gives you is not information. It is a way of working, watched up close, and corrected in real time when you get it wrong. Documentation cannot correct you. A tutorial cannot see that you skipped a step. ## How to find teachers when nobody assigns you one This is the practical part, because most people reading this did not get a mentor handed to them either. What worked for me: - Ask to watch. Most skilled people will let you sit next to them if you are quiet and useful. Sweep the floor, carry the boxes, and pay attention. - Bring finished work, not questions. "I built this, what would you change" gets a real answer. "How do I start" gets a shrug. - Pick people who are good at the craft, not people who are good at talking about the craft. The second kind is easier to find and teaches you less. - Stay long enough to be corrected. The first week of any apprenticeship is politeness. The correction starts when they trust you can take it. ## Reviewing code is being taught Later, when I started reviewing and auditing systems for other people, I realised I was still learning the same way. Every codebase I open was written by someone who made decisions I would not have made, and about a third of the time they were right and I was wrong. That third is where I learn. Over 600+ projects it adds up to an education that no course could have sold me, because no course has seen that many ways for a system to be built and to fail. If you are self-taught, the code review is your classroom. Read other people's systems. Read the ones that work and the ones that broke. Ask why the decision was made before you decide it was wrong. ## What I owe I do not say self-taught anymore without adding the names. I learned in a print shop, at a repair bench, on construction sites, in a foster home with eight children where you learn fast that nobody is going to do your part for you. I learned from people who got nothing for teaching me except the fact that I showed up the next day. If you are the one who learned that way, find the people who taught you and tell them. And then be one of them for somebody else. Not a course, not a thread of advice. Sit next to someone, watch them work, and correct them when they skip a step. That is what self-taught actually means. It means you were taught by people who were not paid to do it. --- ## Data provenance for scientists, explained by an engineer 2026-04-18 · science, data Take a number from the results table of any paper. Where did it come from? Not 'from the data', but precisely: which raw file, which rows, which filters, which transformations, which version of which script, run by whom, on what date. If you can answer that in under a minute, you have provenance. If the answer involves someone's memory, you have a story. ## Provenance is a chain of custody Forensics has a phrase for this: chain of custody. Every time evidence changes hands, someone signs. If the chain breaks, the evidence is inadmissible, no matter how convincing it looks. Data deserves the same treatment. The claim at the end is only as strong as the weakest link between it and the observation. In software we build this by never mutating in place. Raw data lands once and is frozen. Every step that derives something writes a new artifact and records what it read, what code ran, and what it produced. The result is a graph: nodes are datasets, edges are transformations. Walk it backwards from any number and you arrive at the instrument that measured it. ## The three things to record - What: a content hash of every input and output, so you can prove nothing was silently changed - How: the exact code and environment that ran, by commit hash and dependency lock, not by description - Who and when: the person or process that triggered the run, and the timestamp That is it. Three fields, recorded automatically by the pipeline rather than typed into a lab notebook. The automation matters more than the fields. Provenance that depends on people remembering to write it down decays within a month. ## What it looks like in practice You do not need a graph database. A directory convention and a small script get you most of the value. Each derived dataset lives in its own folder with a manifest file: inputs by hash, the git commit, the command line, the environment lock, the run time. The script that produces the dataset also writes the manifest. If the manifest is missing, the dataset is not trusted. If an input hash does not match the manifest, the pipeline refuses to run. When someone asks why the effect size changed between the draft and the submission, you diff two manifests instead of holding a meeting. ## Why edits in place destroy everything The most common provenance failure I have seen in scientific data is the helpful edit. A collaborator notices an obvious typo in a spreadsheet, fixes it, saves the file. The fix was correct. It is also now invisible. Nobody can tell which cells were touched, whether the edit was made before or after the analysis ran, or whether a second, less obvious fix happened in the same session. The cost of that convenience is the entire chain of custody. The rule is simple and non-negotiable: corrections are transformations. They live in code, with a comment explaining why, applied on top of the untouched original. ## Provenance is what makes a tool trustworthy Any system that claims to help with research, including anything built with AI, has to carry provenance through every step or it is producing plausible text, not evidence. The output must be able to say, for every claim, which sources support it and how. I hold my own work to that standard because I have watched what happens to conclusions that cannot be traced: they get repeated, then cited, then believed, and by then nobody remembers they were never checked. --- ## Streaming UI and what it costs the backend 2026-04-15 · fullstack, nextjs Streaming is the best thing that happened to server rendering in a decade. The shell of the page arrives in a hundred milliseconds, the slow parts fill in as they resolve, and the user starts reading before the database has finished thinking. I use it on almost every product I build. But streaming moves cost, it does not remove it. The frontend feels faster because the backend is doing the same work under a different shape, and that shape has consequences nobody puts in the demo. ## The connection stays open A traditional render holds a connection for as long as the slowest query, then sends everything at once. A streamed render holds it for exactly as long, but the first byte goes out immediately. The total connection time per request does not go down. What changes is that the user perceives it as fast. That matters for capacity. If your recommendations panel takes three seconds and you stream it, every page view keeps a connection open for three seconds plus. Your reverse proxy, your load balancer and your serverless runtime all have limits on concurrent open connections, and a slow streamed component turns a throughput problem into a concurrency problem. I have seen a team stream a slow panel to fix a Core Web Vitals score and then hit their platform's concurrent request cap on a Tuesday afternoon. The score went up. The site went down. The panel was the same three seconds either way. ## Suspense boundaries fan out Each Suspense boundary is a promise the server is waiting on. Put five on a page and the server is running five independent data paths for every request. Before streaming, that page would have made one or two queries in sequence. After, it makes five in parallel, and each one holds a database connection from the pool. Parallel is good for latency and bad for pools. A connection pool of twenty with five queries per page view saturates at four concurrent users. I check the pool size against the number of boundaries per route on every audit, and the math is wrong more often than it is right. The defense is the same as everywhere: dedupe the reads, cache what does not change, and give each boundary a query that is fast on its own. A streamed component that runs an unindexed query is still an unindexed query. It just has a spinner now. ## Slow paths hide better The most subtle cost is observability. When a page rendered as one unit, a slow query made the whole page slow and someone noticed. When it streams, the shell is fast, the metric that everyone watches looks fine, and the three-second panel becomes invisible in the dashboards because time to first byte no longer includes it. So I instrument per boundary. Each async component reports its own duration with the route and the boundary name attached. The alert is on the slowest boundary, not on the page. Otherwise streaming becomes a way to hide regressions behind a skeleton. Errors need the same treatment. A boundary that throws renders its error fallback and the rest of the page is fine, which is the right behavior for the user and the wrong behavior for the on-call engineer if nobody logged it. Every error fallback reports before it renders. ## Where streaming earns its cost Streaming is worth it when a page has one or two slow, non-critical parts and a shell the user can act on immediately. A product page with a fast header and a slow reviews section. A dashboard where the summary is cached and the detail is live. It is not worth it when every part of the page is slow, because then you are just paying the connection cost to show a skeleton for three seconds. Fix the queries first. And it is not worth it when the slow part is the thing the user came for, because a fast shell around a slow answer is a fast way to disappoint. Design the budget per route: how many boundaries, how long the slowest one may take, how many connections that implies at peak. Then stream. The feature is excellent. The invoice still arrives. --- ## When not to use a language model 2026-04-11 · ai, system-design I ship language models in production and I like them. That is exactly why I spend so much time telling teams not to use one. A model is a component: expensive per call, slow compared to a function, non-deterministic, and wrong some percentage of the time in ways you cannot fully predict. Those are fine properties for some jobs and disqualifying for others. The skill is knowing which is which before the feature is built, not after the incident. ## The question I ask first Before any AI feature, I ask: can this be computed? If the answer is a lookup, a rule, a calculation or a query over data you already have, a model is the wrong tool, not because it cannot do it but because something else does it perfectly, instantly, for free, and with a stack trace when it fails. A model that decides whether an order qualifies for free shipping is a liability. A function that does it is a fact. The model earns its place only where the input is unstructured, the rules are fuzzy, or the output is language. ## Cases where a model is the wrong tool - Anything with a correct answer that must be exactly right every time: prices, totals, dates, legal thresholds, dosages, IDs. - Anything on a hot path with a tight latency budget where a few hundred milliseconds is a regression. - Anything that runs at a volume where per-call cost dominates, like classifying every log line or every event. - Anything where you cannot build an eval set because nobody can define what a good output is. - Anything where the user needs to trust the answer and cannot check it, unless a human is in the loop. ## The hybrid that usually wins The best designs I have seen use the model at the edges and code in the middle. The model turns messy input into a structured request: extract the fields, classify the intent, normalize the entity. Deterministic code then does the actual work: validate, compute, look up, decide. The model may come back at the end to phrase the result for a person. This keeps the non-determinism in the two places where it is harmless, understanding and phrasing, and out of the place where it is dangerous, the decision. It also makes the feature testable, because the middle is ordinary code with ordinary tests. ## The cost of being wrong The decision is not about whether the model can do the task. Frontier models can do almost anything on a demo. It is about the cost of the times it will not. A summarizer that is wrong one time in fifty is a great product. A dosage calculator that is wrong one time in fifty is a lawsuit. The same error rate, the same model, opposite verdicts. When I evaluate a proposal, I ask for the error rate they expect and the consequence of a single error, and I multiply. If the product is not survivable, the model does not go in the decision path, whatever the demo looked like. ## A model is a component The mindset that makes this easy is to treat the model like a database or a queue: a piece of infrastructure with known characteristics that you place where those characteristics fit. Nobody puts a message queue in the middle of a synchronous checkout because queues are exciting. Nobody should put a language model in the middle of a deterministic rule because models are exciting either. Put it where language is, keep it out of where arithmetic is, and most of the hard decisions make themselves. --- ## Designing for the slow dependency, not the dead one 2026-04-09 · system-design, reliability Everyone designs for the dependency that dies. The connection is refused, the call fails in milliseconds, the circuit breaker opens, the fallback kicks in. It is clean. It is also the rare case. What actually takes systems down is the dependency that gets slow. It still accepts connections. It still answers, eventually. It just takes two seconds instead of twenty milliseconds, and every one of those two seconds is a thread, a connection and a memory buffer that your service is holding open on its behalf. ## How slow becomes dead Say your API handles 200 requests per second and calls a recommendation service that normally answers in 20 ms. You hold around 4 connections to it in flight on average. Now that service degrades to 2 seconds. In-flight climbs to 400. Your connection pool is 50. Requests queue behind the pool. Your HTTP worker threads fill up waiting for the pool. Now your health check, which never touched recommendations, cannot get a thread and fails. The load balancer pulls you out. You are down, and the thing that killed you is still returning 200. This is why the circuit breaker alone is not enough. Most breakers trip on errors. A slow dependency does not produce errors until timeouts fire, and if your timeouts are 30 seconds, the pool is long gone. ## Deadlines, not just timeouts Every outgoing call needs a timeout that is tight relative to normal latency, not relative to the worst case you can imagine. If p99 is 80 ms, a 500 ms timeout is a real defense. A 30 second timeout is a formality. Better than a timeout is a deadline that propagates. The incoming request has a budget, say 800 ms. Each downstream call receives what remains. If the first call ate 600 ms, the second one gets 200, and if that is not enough, fail now instead of doing work nobody will wait for. Most RPC frameworks and gRPC support this natively; over plain HTTP you carry it in a header. ## Bulkheads: one pool per dependency The failure above happened because one pool served everything. Separate them. One connection pool, one thread pool or semaphore, per dependency, sized so that saturating it cannot exhaust the service. If recommendations get at most 20 concurrent calls, then when they hang, 20 requests suffer and the other 180 per second keep flowing. The name comes from ship compartments: one floods, the ship stays afloat. Bulkheads are also where load shedding lives. When the semaphore is full, do not queue. Reject immediately with a fallback. Queuing behind a slow dependency is just a slower way of falling over. ## Fallbacks with last-known values For most non-critical data, a stale answer beats no answer. Cache the last successful response per key with a timestamp, and when the call times out or the bulkhead is full, serve the cached value and mark it as stale in your metrics. Recommendations from ten minutes ago are fine. Exchange rates from ten minutes ago are usually fine. A patient's allergy list is not, and that distinction is a product decision you make explicitly, not one you leave to the cache. For the dependency that really is critical, the fallback is a fast, honest error. A 503 in 100 ms is kinder to the user and to your system than a spinner for 30 seconds. ## Hedging and what to measure For reads that are idempotent, hedged requests help: send the call, and if it has not answered by the p95 latency, send a second copy to another instance and take whichever returns first. It costs a few percent of extra load and cuts tail latency dramatically. Never hedge writes. Finally, measure p99, not average. An average of 40 ms can hide one percent of calls taking 4 seconds, and that one percent is exactly the traffic that fills your pools. Alert on tail latency and on pool saturation per dependency. Those two graphs tell you a dependency is getting slow long before it takes you with it. --- ## Structured logs or nothing 2026-04-05 · infra, observability There are two kinds of logging in the systems I audit. The first kind is text: 'User 4412 checked out, total 189.90, took 340ms'. The second kind is data: a JSON object with user_id, event, total_cents, duration_ms and a dozen context fields. Both contain the same information. Only one of them is useful at 3 a.m. with a customer on the phone. The first kind requires a human to read it. The second kind can be filtered, grouped, counted, joined and alerted on by a machine. Once you have worked with the second kind, the first one feels like writing a database as a novel. ## What structured actually means Structured logging is not 'we output JSON'. It is a contract about fields. Every line has the same base set: timestamp, level, service, version, environment, a request or correlation id, and a short event name that identifies what happened. On top of that base, each event adds its own fields, named consistently across the codebase. The event name is the field people forget. 'message' is free text and will be different every time someone touches the code. 'event' is an identifier like order.paid or auth.login_failed. It is what you group by, what you alert on, and what you search for when the free text has been reworded three times. Consistency of names is the other half. If one module logs user_id and another logs userId and a third logs uid, the query that joins them will be written wrong, and the investigation will conclude that the user was never seen. Decide the field names once, put them in a shared logger module, and have the logger reject unknown top-level fields if the platform allows it. ## The logger is one module In a TypeScript codebase I insist that logging goes through exactly one module that wraps whatever library the team chose. The module owns the base fields, attaches the correlation id from async context so nobody passes it by hand, redacts known sensitive keys before output, and exposes a small API: a child logger with bound context, and level methods that take an event name and a fields object. That single point of control is what makes the rest possible. When the observability platform changes, one file changes. When a new field must be redacted, one file changes. When a request id needs to flow from a Next.js middleware through a server action into a database call, the module attaches it from async local storage and the rest of the code never knows. ## What to log and what to keep out Log every state transition that matters to the business, with the identifiers needed to find it again. Log every call to an external system with its duration and outcome. Log every decision that a human might later question: why a request was rate limited, why a payment was retried, why a job was skipped. Keep out anything that is a secret, a full request body, a password field, a token, a card number, or personal data that is not needed to find the record. Redact at the logger, not at the call site, because call sites are written in a hurry. Keep out the noise in hot loops: sample it, or aggregate it into a counter, but do not emit a line per iteration. ## Rules I apply in reviews - Every log line has an event name that is a stable identifier, not a sentence. - Every log line in a request carries the correlation id automatically. - Field names are consistent across the codebase and documented in one place. - Sensitive fields are redacted by the logger, and the list is tested. - Levels mean something: error is 'a human should look', warn is 'expected but unusual', info is 'the business timeline', debug is off in production. None of this is expensive. It is a day to set up the module and a habit to maintain. The return is that the next incident is investigated with queries instead of scrolling, and the customer support question 'what happened to my order' has an answer in the time it takes to type the id. --- ## The modular monolith in practice 2026-04-01 · architecture The modular monolith is the architecture I recommend most often and the one I see done wrong most often. The idea is simple: one deployable, internally split into modules with explicit boundaries, so you get the design discipline of services without the network. The failure is also simple: the boundaries exist in a diagram and nowhere else. This is what it looks like when it works, in the kind of Next.js and TypeScript codebase I spend most of my time in. ## The shape One repository, one deploy unit. Inside it, a directory per module, named after a business capability, not a technical layer: billing, catalog, scheduling, identity. Not controllers, services, repositories. Layers can exist inside a module; they do not organize the top level. Each module has a public surface, a single index file that exports what other modules may use: a handful of functions, some types, maybe an event type. Everything else in the module is private. Other modules import from the index and nothing deeper. Each module owns its data. Its tables, its migrations, its queries. No other module writes to those tables, and ideally no other module reads them directly. If catalog needs to know whether a customer is active, it asks identity through the public surface, not with a join. Modules talk in two ways: direct calls through the public surface, for things that must be synchronous and consistent, and in-process events, for things that are reactions. When an order is paid, billing emits an event and fulfillment subscribes. Same process, same transaction boundary if you want it, no broker. ## How the boundary is enforced A boundary that is not enforced by tooling is a suggestion. The enforcement I set up is cheap. - A lint rule that forbids imports from another module's internals, with the index file as the only allowed entry point. - A separate database schema or table prefix per module, and a CI check that no migration touches another module's tables. - A dependency graph test that fails the build when a cycle appears between modules. - A test that each module's public index exports only what a short allowlist says it should. Those four rules take a day to set up and catch most of the cheating I would otherwise find in an audit two years later. ## Where teams cheat The first cheat is the shared module. It starts as types and helpers, becomes the place where every rule that touches two modules goes, and ends as the real monolith with the others as thin wrappers. The fix is to treat shared as a strict library: pure, no I/O, no business logic. The second cheat is the join. Under deadline, someone writes a query that reads two modules' tables in one statement because it is faster than two calls. It is faster. It also means those modules can never be separated, and the query is invisible to both owners. If the join is truly needed for performance, it belongs in a read model that one module explicitly owns and populates. The third cheat is the transaction that spans modules. A request handler opens a transaction, calls billing, then calls scheduling, then commits. It works and it is convenient. It also means the two modules share a failure mode and a lock scope, and when one is extracted the atomicity silently disappears. My rule: a transaction belongs to one module. Cross-module effects happen through events, with the same at-least-once discipline you would use across a network. The fourth cheat is the framework leak. In Next.js it is the server action or route handler that does the business logic inline, with the module as an afterthought. The handler should be ten lines: parse, call the module, format. If it is a hundred, the module boundary is fiction. ## Why this is worth the discipline A well-kept modular monolith gives you almost everything people extract services for: clear ownership, contracts between parts, the ability to reason about one capability in isolation. It gives it without a network, without a distributed transaction, without a fleet of pipelines. And it keeps the option open. When one module truly needs to scale or deploy on its own, the extraction is mechanical: the public surface becomes an API, the in-process events become a topic, the tables move. I have watched teams do that in weeks because the boundary already existed, and I have watched other teams spend a year because it did not. Start with one deploy unit and hard boundaries. Enforce them with tooling, not intentions. Then extract only what has earned it. --- ## The hidden cost of every dependency you add 2026-03-31 · fullstack, audit A dependency costs nothing to install and something every month forever. That second part is the one nobody budgets for, and it is why a two-year-old project has 1,400 packages in its lock file and nobody can explain what a third of them do. I audit dependency trees the way I audit databases: by reading what is actually there, not what the team thinks is there. The results are the same every time. A handful of packages doing real work, and hundreds pulled in by a utility that saved someone twenty minutes. ## What you are actually buying When you add a package, you buy its transitive tree. A single date library can bring in forty packages. Each of those has a maintainer, a release schedule, a set of open issues and a chance of being compromised. Your project now depends on all of them, and your security scanner will remind you of that weekly. You buy its upgrade cadence. A package that releases a breaking major version every year will cost you a migration every year, and a package that has not released in three years will cost you a fork the day it breaks on a new Node.js version. You buy its bundle weight, if it runs in the browser. With Server Components the calculus changed: a heavy package imported only on the server costs nothing to the user, and the same package imported from a Client Component costs every visitor a download. I check which side of the boundary each dependency sits on, because moving one import across it can be worth more than any amount of tree shaking. And you buy its types. A package with hand-written, out-of-date type definitions will lie to your compiler, and the lie will surface as a runtime error in the one place you trusted the type. ## The questions before adding one I ask a fixed set of questions before a new dependency enters a project I am responsible for, and I ask them out loud so the answer is in the pull request. - Could thirty lines of our own code do this, and would we understand those lines in two years? - How many packages does it bring in, and how many of them have a single maintainer? - When was the last release, and how many open issues mention a crash or a security problem? - Does it ship its own types, and are they generated or hand-written? - Does it run on the server, the client or both, and does it need Node.js APIs that break at the edge? - What is the license, and does our product's license allow it? Most packages fail on the first question. The ones that pass are usually doing something genuinely hard: cryptography, date and time zone math, parsing a format with a spec. Those are worth their cost. A helper that pads a string is not. ## Keeping the tree honest The lock file is committed, always, and the install in CI uses the frozen mode that fails if the lock file and the manifest disagree. That single setting has prevented more surprise breakages than any other configuration I know. Versions are pinned to a range that allows patches and rejects majors. Automated upgrade pull requests run weekly, one package per PR, so that the one that breaks the build is identifiable. A bulk upgrade of forty packages that fails tells you nothing. Once a quarter, I run a report of packages that are installed and not imported anywhere, and remove them. In most projects that list is not short. And I run the audit tool, but I read the results rather than silencing them, because a vulnerability in a dev-only build tool and one in the JWT library are not the same severity, whatever the scanner says. ## The dependency you forgot you have The most expensive dependency in most projects is not in the manifest. It is the hosting platform's proprietary API, the payment provider's SDK that shapes your data model, the auth service that owns your user ids. Those do not show up in a package count and they are the hardest to leave. Treat them the same way. Wrap them behind an interface you own, keep the surface small, and know what it would cost to replace them before you need to. The install is free. The exit never is. --- ## The ethics of building tools for medicine when you are not a doctor 2026-03-28 · science, life I am not a doctor. I studied three years of Biomedicine and left. My father practiced medicine for more than forty years and I met him once. I have spent my career building and auditing software, and now some of that software points at the cardiovascular field. I think about the ethics of this constantly, not as an abstract question, but as a set of concrete rules I refuse to break. ## The tool does not decide The first rule is that the tool never makes the decision. It can search, structure, surface, compare, compute, flag. It cannot conclude on behalf of the person responsible. This is not modesty. It is a design constraint with consequences: every output must be phrased as an input to a human judgment, and the interface must make that judgment explicit. Who looked at this, when, and what did they decide. Software that hides the human step is not more efficient. It is moving accountability from a person to a system that cannot carry it. ## Know exactly what you do not know I know how data breaks in transit. I know how timestamps lie, how identifiers collide, how a null hides a process. I do not know what a borderline measurement means for a specific patient, and I should not pretend that reading papers makes me know. The ethical move is to be precise about the boundary. My work stops at the point where the data becomes a clinical or scientific judgment, and at that point a clinician or a scientist takes over, with full visibility into what the system did and did not do. This is why I build with clinicians and scientists, not for them. Not as advisors who review the finished product, but as the people who define what correct means before I write a line. ## Uncertainty is not a UX problem Product teams are trained to reduce friction and hide complexity. In medicine and science that training is dangerous. A confidence interval is not clutter. A missing data warning is not a distraction. If the system is unsure, the person using it must see that it is unsure, in a way that cannot be dismissed with a click. I would rather ship an interface that looks cautious than one that looks confident and is wrong. ## Speed cannot buy rigor Everything about my field rewards speed. Ship, measure, iterate. That is fine when a mistake is a rollback. In science, a plausible but unvalidated result can propagate for years. So the rule is that any advance has to go through scientific validation before it is treated as an advance, even when the validation is slow and the demo was impressive. This is the one place where I am willing to be the slow person in the room. ## Why I do it anyway If the risks are real, why build these tools at all? Because the alternative is not that nobody builds them. It is that they get built by people who have never audited a production failure at 3 a.m. and who think a green dashboard means a working system. I have seen what happens when software with no discipline meets a domain with no tolerance for error. I would rather be in the room, bringing the discipline, than outside it complaining. My faith teaches me that a life is not a data point, and my career taught me that data points get lost all the time. Building for medicine as an engineer means holding both of those at once. The patients my father treated for free never saw a dashboard. The tools I build have to be worthy of people like them. --- ## When to add a message broker, and when you are avoiding a decision 2026-03-25 · system-design Someone on the team says: we should put a queue between these two services. It sounds like architecture. Sometimes it is. Often it's a way to postpone a question nobody wants to answer. I have reviewed enough systems to recognise the difference, and I want to lay out the tells. ## What a broker actually buys you A message broker gives you four concrete things. Decoupling in time: the producer can write while the consumer is down, and the consumer catches up later. Buffering: a spike of ten thousand orders becomes a backlog instead of a wave of 503s. Fan-out: one event, five consumers, none of which the producer needs to know about. Retries with backoff and a dead letter destination, handled by infrastructure instead of by hand-rolled loops. If you need at least two of those, and you can name them, a broker is probably right. If you can only say 'it's more scalable', you don't have a reason yet. ## What it costs Every one of those benefits has a bill. Ordering becomes conditional: most brokers guarantee order per partition or per key, not globally, and the moment you add a second consumer you gain parallelism and lose sequence. Duplicates become normal, so every consumer has to be idempotent, which is a whole discipline of its own. Visibility drops: a synchronous call fails in your face, an asynchronous one fails in a dead letter queue you have to remember to watch. The broker is also another system to run, with its own upgrades, disk, partitions, consumer group rebalances and 2 a.m. pages. And the messages are a contract. Once three consumers depend on the shape of order.created, changing a field is a versioning project, not an edit. ## The tell: a queue instead of a decision Here's the pattern I see most. Two teams cannot agree on who owns a piece of data or a transaction. Instead of deciding, they put an event between them. Service A publishes 'thing happened', service B consumes it and updates its own copy. Now nobody owns the truth, the two copies drift, and the queue is where the disagreement goes to hide. The other version: an operation that should be one transaction gets split into 'save the order' and 'publish the event', with no outbox, so the event is sometimes lost and sometimes published for an order that rolled back. The broker did not create that inconsistency. It just made it invisible. Ask directly: what is the source of truth for this data, and where is the transaction boundary? If the queue is the answer to either question, it's not an answer, it's a deferral. ## Start with a table For a large share of cases, a jobs table in Postgres is the right first broker. Insert a row in the same transaction as the business change, so the job can never exist without its data or the data without its job. Workers poll with SELECT ... FOR UPDATE SKIP LOCKED, which lets many workers pull from the same table without stepping on each other. Retries are an attempt counter, dead letters are a status column, visibility is a query, ordering is an index. There is nothing new to operate. This handles thousands of jobs per minute comfortably on ordinary hardware, and libraries exist for every major stack. Most systems I audit never outgrow it, and the ones that do understand their requirements far better by the time they leave. ## When the real thing is right - The volume is genuinely high, in the tens of thousands of messages per second. - You need fan-out to many independent consumers that will change over time. - Producers and consumers belong to different teams with different deploy cycles. - You need replay of history, which log-style brokers do well and tables do badly. - The workload is a stream, not a set of jobs, and you want windowing and stream processing. Even then, keep the outbox pattern on the producer side. The broker doesn't solve atomicity between your database and the log. Only a transaction does. The question is never 'should we use a queue'. It's 'what does this queue let us stop deciding', and whether that's a decision we're allowed to postpone. --- ## Prompt injection is input validation you forgot 2026-03-20 · ai, security Every generation of engineers relearns the same lesson: untrusted input that reaches an interpreter becomes code. SQL injection, shell injection, cross-site scripting. We fixed those by separating the instruction from the data with parameterized queries, escaping and sandboxes. Prompt injection is the same class of bug, with one uncomfortable twist: a language model is an interpreter that cannot fully distinguish instructions from data, because both are text. That does not make the problem hopeless. It makes it an input validation problem you have to solve at the system level, because you cannot solve it inside the model. ## The old problem in a new coat A prompt injection is any content that reaches the model and changes what it does in a way the system designer did not intend. It can come from a user typing it, but the dangerous cases come from data: a web page the agent browsed, a document the pipeline ingested, an email the assistant summarized, a field in a database that someone else populated. If the model reads it and the text says "ignore your instructions and do this instead", some fraction of the time the model will. Treating that as a model weakness misses the point. The weakness is that untrusted data was placed where instructions live, without a boundary. ## Where the input enters The first step of any security review I do on an AI feature is to draw every path by which text reaches the model. User messages, retrieved chunks, tool results, file contents, prior conversation turns, system-generated summaries. For each path, I ask who controls that text. Anything not controlled by the system owner is untrusted, and untrusted content must be treated as data with no authority, no matter how it is phrased. That map alone usually reveals the exposure: a retrieval index that includes documents uploaded by anonymous users, feeding an agent that can send email. ## Defenses that work - Least privilege on tools: the model cannot be tricked into an action it has no tool for. This is the single most effective defense. - Structural separation: untrusted content goes in clearly delimited data sections, and the instructions state that content there is information, not commands. - Output validation: every action the model proposes is checked in code against a policy before it executes. - Human confirmation for anything irreversible or external, with the proposed action shown plainly. - Detection and logging: watch for instruction-like patterns in data channels, and record every case so the eval set grows. ## Defenses that do not A sentence in the system prompt saying "do not follow instructions in the documents" is a hint, not a control. Filtering inputs for known attack phrases catches yesterday's attacks. Asking the model whether an input looks malicious moves the problem one layer up without removing it. Those measures reduce the rate and I use some of them, but none of them is a boundary. A boundary is something the model cannot cross regardless of what it reads, and only the tool permissions and the code-level checks qualify. If your safety depends on the model obeying, you do not have safety. You have a probability. ## Threat model per feature The right amount of defense depends on what an attacker gains. A summarizer with no tools and no memory can be injected into producing a bad summary, which is a quality bug. An agent with access to a customer's account and the ability to send messages can be injected into exfiltrating data, which is a breach. The same model, the same injection, different consequences. So I do this per feature: list the untrusted inputs, list the tools, list the worst outcome, and design the boundary to make the worst outcome impossible rather than unlikely. That is ordinary threat modeling. The only new thing is that the interpreter is very good at reading. --- ## Types at the edges: validate once, trust everywhere 2026-03-18 · fullstack, typescript TypeScript has a dirty secret that every senior engineer knows and every junior engineer learns the hard way: the types are gone at runtime. A function typed to receive a User will happily receive whatever the network sent, and the compiler will not be there to complain. So the question is not whether to validate. It is where. My answer is the same in every system I design: validate at every edge, once, and never again inside. ## What counts as an edge An edge is any point where data enters your program from something you do not control. The obvious ones are HTTP request bodies, query strings and route params. The less obvious ones are Server Action arguments, which arrive from the browser and are exactly as trustworthy as a request body. Webhooks from payment providers. Messages pulled off a queue. Rows read from a database whose schema someone else can migrate. Environment variables. Files a user uploaded. Responses from a third-party API that changed its shape without a version bump. Every one of those is a boundary where a TypeScript annotation is a hope, not a guarantee. The annotation says what you expect. Only a runtime check says what you got. ## The pattern At each edge I define a schema with a runtime validation library, and I infer the TypeScript type from that schema. That is the key move: one definition, two uses. The schema validates at runtime and the type flows through the compiler. They cannot disagree because they are the same object. The parsed value is the only thing that crosses into the interior. The raw input never does. If a handler receives unknown, calls parse, and passes the result down, every function below it can trust its parameter types completely. There is no second validation layer, no defensive null checks in the middle of business logic, no 'just in case' coercion. That interior trust is the payoff. It is what makes the code readable. A pricing function that receives a validated Order can spend its lines on pricing, not on wondering whether quantity is a string. ## The details that matter Parse, do not just validate. A validator that returns a boolean leaves you with the original value and a promise. A parser returns a new value of the narrowed type, with defaults applied and unknown fields stripped. Stripping unknown fields is a security property: it stops a client from smuggling an isAdmin flag into an update. Fail loudly at the edge, softly nowhere. A malformed request gets a 400 with a structured error. A malformed webhook gets logged with the raw payload and rejected. A malformed environment variable crashes the process at startup, before it serves a single request, which is the cheapest possible time to discover it. Validate what you send, too. When your server responds to a client, or publishes an event, that output is someone else's edge. Checking it against a schema before it leaves catches the bug where a refactor changed a field name and every consumer broke at once. Branded types close the loop. Once an email string has passed the email schema, give it a branded type so that a function requiring a validated email cannot be called with a raw string. The compiler now tracks which values have been checked, and the proof travels with the value. ## What I refuse to do I do not validate twice. Double validation is a sign that nobody trusts the first check, and it means the interior is full of redundant guards that hide real logic. I do not use type assertions at the edge. Writing `as User` on a fetch result is a lie to the compiler, and the compiler believes you. Every audit where I found data corruption in a database, I found an assertion somewhere upstream. And I do not skip the edge because 'it is our own API'. Your own API is built by a person who will change it on a Tuesday. The schema check is how you find out on Tuesday instead of a month later, from a customer. --- ## The first thing I check in any codebase 2026-03-15 · audit, engineering When a codebase lands on my desk, I do not start with the README, the architecture diagram or the test suite. I start with a search: every place the system writes to its database. Inserts, updates, deletes, upserts, ORM saves, raw SQL, stored procedures. I list them all, and then I count. That number tells me more about the system in ten minutes than a day of reading. ## Why the write path Reads can be wrong and the damage is temporary: a stale page, a wrong number on a screen, refreshed away. Writes are permanent. A wrong write is a corrupted row that every future read will faithfully return, forever, until someone notices. Every serious data bug I have found across 600+ projects was a write that should not have happened, happened twice, happened halfway, or happened to the wrong row. If I only have an hour, the write path is where the hour goes. ## What the count tells me If there are five places that write to the orders table, I know that business rules for orders live in five places and probably disagree. If there are fifty, the team has no concept of a write boundary and every feature has invented its own. If there is one, someone thought about this, and I go check whether the one place is actually used or whether the other forty-nine bypass it through a raw query 'just for the migration script'. The count is a proxy for discipline. Low and centralized means the system has a spine. High and scattered means the data model is whatever the last feature needed it to be. ## The three questions at every write For each write site I ask the same three things. Is it inside a transaction with the other writes it logically belongs with, or can the system crash between them and leave half a state? Is it idempotent, meaning that running it twice with the same input produces the same result, or does a retry create a duplicate? Is it validated at this boundary, or does it trust that whoever called it already checked? Most write sites fail at least one. The ones that fail all three are where I start the report. ## What I find, reliably An update that sets a status without checking the previous status, so a cancelled order can become paid. An insert inside a loop with no transaction, so a crash at item seven leaves six orphans. A delete that runs before the thing that depended on it was updated. A save that is called from an event handler and again from the API route that raised the event. None of this is exotic. It is what happens when writes are scattered and nobody owns the path. Only after the write map do I look at architecture, tests and documentation, and now I read them differently. When the diagram shows a clean service boundary, I already know whether the writes respect it. When the tests are green, I already know which write sites have no test at all. The rest of the audit is largely explaining what the write map already revealed. ## Do it on your own code You do not need an auditor for this. Search for every write to your most important table, put the results in a list, and answer the three questions for each. It takes an afternoon. In most codebases it produces at least one finding that would otherwise have shown up as a support ticket with the word 'duplicate' in it, six months from now, at the worst possible time. --- ## The cloud cost nobody budgets for 2026-03-13 · infra, business When a founder asks me to estimate hosting costs, they are thinking about servers. How many, how big, how much per month. That number is usually the easiest to predict and the least likely to surprise anyone. The surprises live elsewhere, in categories that scale with behaviors nobody modeled, and they show up on the invoice months after the decision that caused them. ## Egress, the tax on success Data leaving the cloud provider's network is billed per gigabyte, and the rate is high enough that for media-heavy products it becomes the largest line on the bill. Every image served from object storage without a CDN in front of it, every API response to a mobile client, every backup copied to another provider, every analytics export downloaded by the data team. All of it is egress. The insidious part is that egress grows with usage, not with infrastructure. You can hold servers constant and watch the bill double because customers uploaded more photos and shared more links. A CDN with a generous free tier or cheap egress in front of your storage is not an optimization. It is a requirement, and it should be in the design before the first file is stored. ## Logs and metrics, the cost of curiosity Observability platforms bill by volume ingested and by retention. A single debug log line inside a hot loop can produce terabytes per month. A high-cardinality metric, one with a user id or request id as a label, can multiply the time series count by millions and turn a modest metrics bill into the second largest item on the invoice. I have reviewed systems where the logging bill exceeded the database bill. The fix is almost never 'log less'. It is sampling on the noisy paths, structured fields instead of repeated text, aggressive retention tiers where the last seven days are searchable and the rest is cold, and a rule that no label on a metric can hold an unbounded value. ## NAT gateways, load balancers and the per-hour tax Managed networking components are billed per hour of existence plus per gigabyte processed. A NAT gateway in every availability zone, a load balancer per environment, a VPN endpoint that someone set up for a demo, all of them tick away at a few dollars a day whether or not traffic flows. On a small bill, these fixed costs are often a third of the total and they are invisible because no single one looks expensive. The related trap is cross-zone and cross-region traffic. Two services that talk constantly, placed in different zones for resilience, pay per gigabyte for every conversation. That is often the right trade, but it should be a trade someone made, not one they discovered. ## The things I ask teams to model - Egress per active user per month, including media, API responses and third-party integrations. - Log and metric volume per request, multiplied by expected requests, multiplied by retention days. - Fixed hourly components per environment: gateways, balancers, endpoints, idle databases. - Storage growth per month including backups, snapshots and versioned objects that are never deleted. - Managed service minimums: the smallest instance of a managed database or cache still costs money at zero traffic. Model these per unit of business activity, not per server. 'Cost per thousand active users' is a number a founder can plan with. 'Cost per instance' is a number that hides everything above. ## The habit that prevents surprises Read the bill every month with the line items expanded. Not the total, the lines. Sort by cost, look at the top ten, and for each one ask what behavior drives it and whether that behavior is growing. Set a budget alert at a threshold that would hurt, and a second one at half of that. Cloud pricing is not a trap. It is a mirror of what your system actually does, at a level of detail nobody looked at during design. Reading it carefully is the cheapest audit you will ever run. --- ## Reversibility as an architecture goal 2026-03-11 · architecture, system-design I do not try to make the right decision on every architecture question. I try to make the decision that is cheapest to reverse when it turns out to be wrong. Those are different goals, and the second one has made me far more money for clients than the first. Being right requires information you do not have yet: how the product will evolve, how many users will show up, which integrations will matter. Being reversible only requires discipline you can apply today. And once you can reverse cheaply, being wrong stops being expensive, which means you can afford to decide faster and with less debate. ## Two kinds of decisions Some decisions are doors that swing both ways. Which logging library. How to structure a folder. Whether a job runs every five minutes or every ten. If you get them wrong, you change them, and nobody outside the team notices. Other decisions are one-way doors. The primary database engine after there is data in it. The tenancy model. The identifier format exposed in public URLs. The event schema that three external partners have integrated against. The programming language of the core. Getting these wrong costs a migration project measured in quarters. The mistake most teams make is treating both kinds with the same amount of care. They debate the logging library for a week and pick the tenancy model in an afternoon because 'we can always change it'. The right allocation is the reverse. Spend the meeting time on the one-way doors. Decide the two-way doors in ten minutes and move on. ## Turning one-way doors into two-way doors This is the actual architecture work: taking a decision that would be expensive to reverse and wrapping it so that reversal becomes possible. Not free, but possible, within a sprint or two. The database engine becomes reversible when no module outside the data layer knows which engine it is: no raw SQL leaking into business code, no engine-specific types in the domain, and a test suite that runs against a second engine occasionally, even if only a subset. The public identifier format becomes reversible when it is opaque to consumers from day one, so you can switch from integers to ULIDs behind a mapping table without anyone noticing. The message broker becomes reversible when producers and consumers talk to a thin internal interface with a real second implementation, even if it is just in-memory for tests. None of these are exotic. They are the same techniques people call 'clean architecture' and then abandon under deadline pressure. The framing that keeps them alive is the question: 'if this choice is wrong, what is the cost of finding out?' ## The reversibility budget Reversibility is not free, and pretending otherwise is how teams end up with abstraction layers over everything. Every seam you add costs some indirection, some testing, some cognitive load. So I spend it deliberately. I ask three questions about each decision. How likely is it that we learn something in the next year that would change this? How expensive is a reversal if we do nothing to prepare? And how expensive is the preparation itself? If the likelihood is low and the reversal cost is low, do nothing. If the reversal cost is high and the preparation is cheap, prepare. If both are high, that is the decision that deserves a week of real analysis, a written ADR and probably a prototype. Vendor lock-in gets its own line. Any managed service that stores your data or defines your workflow gets a documented export path before the first production write. Not a full migration plan, just proof that the data can leave. I have seen teams held hostage by a pricing change because they never checked. ## What this looks like in a review When I review a system, I list its one-way doors on a single page. Usually there are fewer than ten. For each one I write down whether it is still reversible, at what cost, and what would trigger the reversal. That page tells the leadership team more about their real risk than any dependency diagram. The systems that age well are not the ones that made every call correctly. They are the ones where the wrong calls could be walked back before they became load-bearing. Design for that. Being right is a bonus. --- ## Building world-class systems from a small city in Paraná 2026-03-08 · career, life I moved to Toledo, in the west of Paraná, in 2018. If you have not heard of it, that is the point. It is a small city surrounded by farms, a long way from the neighbourhoods in São Paulo where I grew up and a longer way from anywhere people imagine software being built. I have designed and audited systems for clients all over the world from here. This is not a post about how location does not matter. Location matters. It is a post about what you have to do differently when you are not in the room. ## The work must speak without you In a big city you can be a mediocre engineer with a good network. Someone vouches for you, you get the meeting, the meeting goes fine. From a small city, nobody vouches for you. The first thing a potential client sees is the work, and the work has to hold on its own. This pushed me to become obsessive about the artefacts. The audit report has to be readable by someone who never talked to me. The architecture document must survive being forwarded three times. The code review has to explain not only what is wrong but why it matters and what to do instead. Over 600+ projects this turned from a survival tactic into the thing I am known for. People say I find failures nobody else finds. Part of that is skill. Part of it is that for years I could not rely on charm, so I relied on evidence. ## Asynchronous by default, not by preference When your clients are in three time zones and your city has no tech meetups, you learn to work asynchronously or you do not work. I write more than I talk. Decisions go into documents with dates. Questions are batched so that one message covers a day, not twelve. This turned out to be better engineering practice, not just a compromise. A decision written down can be audited later. A conversation cannot. When a system breaks two years after I designed it, the reasoning is still there, and the team that inherited it can read why the boundary was drawn where it was. ## The distractions are gone, and so is the noise There is a real cost to being far from the centre. You miss the hallway conversation where someone mentions the tool that would have saved you a month. You have to work harder to stay current. I read more than I would if I could just ask a friend over coffee. But there is also a real gain, and I did not expect it. The noise is gone. The pressure to adopt the framework everyone in the city is adopting this quarter does not reach here. I choose Next.js and TypeScript because they have held up across many projects, not because a meetup said so. Distance from the hype cycle is a technical advantage if you use the quiet to think. ## What I would tell an engineer in a small city You are not at a disadvantage in the work. You are at a disadvantage in the visibility of the work. So make the work visible on purpose. Write down what you did and why. Publish the decisions you are allowed to publish. Answer clearly and in writing so people can forward you. Be the engineer whose documents are better than everyone else's meetings. And do not move to fix a problem that better artefacts would fix. I have built a career, a studio and now a research project from a city most people cannot place on a map. The farms are still outside the window. The systems run in data centres on other continents. That combination is not a limitation. It is just the shape of the job now. --- ## The prompt is not the spec 2026-03-06 · ai, engineering I keep seeing the same artifact in AI features I review: a long system prompt, lovingly edited over months, that everyone treats as the definition of the feature. Ask what the feature should do in an edge case and someone opens the prompt and reads a sentence from it. That is not a spec. That is an implementation detail that happens to be written in English. ## What a spec does that a prompt cannot A spec answers questions the model never sees. What inputs are valid and what happens to invalid ones. What the output must contain, in what shape, and what is forbidden. What the feature must refuse to do. What the acceptable failure modes are and how they surface to the user. What good looks like, with examples, and what bad looks like, with examples. Who is accountable when it is wrong. A prompt can hint at some of this, but the model is free to ignore it, and the next model will ignore different parts. The spec is the thing you test against. The prompt is one attempt at satisfying it. ## The prompt is an implementation detail Treat the prompt the way you treat a SQL query or a regex: versioned, reviewed, tested, and replaceable. When the provider ships a new model, the prompt that worked is now a hypothesis. If your only definition of correct lives inside it, you have nothing to check the new behavior against. Teams that keep a spec outside the prompt swap models in an afternoon. Teams that do not spend two weeks arguing about whether the new tone is a regression. ## Where the spec actually lives In the features I ship, the spec is three concrete things. First, a schema for the output, enforced in code, not requested politely. Second, an eval set: a few hundred inputs with expected outputs or grading rules, covering the happy path, the boring path and the adversarial path. Third, a short document that states the invariants in plain language: this assistant never quotes a price it did not retrieve, this extractor never returns a date it did not see in the source. The prompt is derived from those three. When the prompt and the spec disagree, the spec wins and the prompt gets fixed. ## How I write the spec for an LLM feature I start with the failures, the same way I design any system. What is the worst plausible output? A support assistant that invents a refund policy. A summarizer that drops the one sentence with the contraindication. A classifier that is confidently wrong on the rare class that matters most. Each of those becomes a rule and a handful of eval cases before I write a single line of prompt. Then I write the happy path, which is easy and which is where most teams start and stop. The spec is done when a colleague could build the feature on a different model, from the spec alone, and pass the same evals. ## The tell The tell that a team has confused the prompt with the spec is simple. Ask them to show you the last regression. If the answer is a diff of the prompt and a Slack thread about vibes, there is no spec. If the answer is a failing eval case with a name, there is. The second team is not smarter. They just decided early that the prompt is code and correctness is something you write down somewhere else. --- ## Back-pressure is a feature, not a bug 2026-03-03 · system-design, reliability A system that says "not now" is healthier than a system that says "yes" to everything and then falls over. That is back-pressure: the ability of a slower part of a system to push the pressure back to a faster part instead of silently absorbing it until it breaks. ## What happens without it Picture a service that accepts HTTP requests and writes to a database. The database slows down. Without back-pressure, the service keeps accepting requests, each one holds a thread and a connection waiting on the database, the thread pool fills, memory grows, and eventually the service stops responding to everything, including health checks. The load balancer marks it dead, the traffic moves to the next instance, and the same thing happens there. That is a cascade, and the root cause was a service being too polite to refuse. The same story happens with queues. A producer writes faster than consumers read, the queue grows, lag becomes hours, and the "async" system is now delivering yesterday's events. Nobody refused anything, so nobody noticed until a customer did. ## The mechanisms Back-pressure shows up in a few concrete shapes. Bounded queues. An unbounded queue is a memory leak with a delay. Give every in-memory queue a maximum size and decide what happens when it is full: block the producer, drop the oldest, or reject the newest. Pool limits with fast failure. When a connection or thread pool is exhausted, fail immediately with a clear error instead of queueing the request behind others that are already stuck. A request that fails in five milliseconds can be retried. A request that waits thirty seconds and then fails has cost you thirty seconds of a thread. Explicit rejection. HTTP 429 and 503 with a Retry-After header. They tell the caller the truth: the system is here, it is busy, come back in a moment. That is far kinder than a timeout. Pull-based consumption. In a queue system, consumers ask for work at the rate they can handle it. Log-based brokers do this naturally; push-based webhooks do not, which is why webhook receivers need their own buffer and their own limits. ## Load shedding is a product decision When you must refuse, which requests do you refuse? This is where engineering meets product. A checkout should keep working while recommendations are shed. Reads can be served stale while writes are protected. Anonymous traffic can be limited before authenticated traffic. Deciding this in advance, and encoding it in priorities or separate pools, is what makes back-pressure a feature. Deciding it during the incident is what makes it a bug. I like to write the priority order down in a sentence: "Under pressure, this system protects payments, then logged-in reads, then everything else." If a team cannot write that sentence, their system will pick the order at random at 3 a.m. ## The signals to watch Queue depth and its growth rate, not just its size. Pool utilisation, because a pool at 95% is about to become a pool at 100%. Time spent waiting for a resource versus time spent using it. The ratio of rejected to accepted requests, which should be visible and alerted on, because a rising rejection rate is the system telling you it is doing its job and needs help. ## The mindset Teams resist back-pressure because a 429 looks like a failure on a dashboard. It is the opposite. It is the one failure mode that is bounded, observable, and recoverable. A system that never says no has no way to protect the requests that matter from the ones that do not. Build the refusal path with the same care as the happy path. Then, when the pressure comes, your system bends instead of breaking. --- ## AI and the end of boilerplate, and what fills the space 2026-02-28 · ai, career I am self-taught, and a large part of my early learning was typing boilerplate. Controllers, DTOs, migrations, the same form validation for the fortieth time. That work taught me patterns by repetition, and it also consumed years. Today a model writes that code in seconds, usually correctly. I do not mourn it. But I pay attention to what fills the space it left, because the space is not empty, and the people who assume it is are the ones who get surprised. ## What actually got cheap The cost that collapsed is the cost of producing plausible code from a clear description. Scaffolding, glue, adapters, tests for obvious behavior, the translation of a known pattern into a specific language. If you can describe it precisely, the model can write it. That "if" is doing a lot of work. The precision of the description is now the bottleneck, and precision is exactly the thing junior engineers used to learn by writing the boilerplate themselves. We have removed the practice ground without removing the need for the skill. ## What fills the space Three things, in my experience. Specification: deciding what the system should do, including the edge cases, before anyone or anything writes code. Verification: reading generated code with the assumption that it is wrong somewhere, and building the tests and evals that prove otherwise. And judgment about structure: knowing which of the five working solutions the model offered is the one that will still be maintainable when the requirements change. None of these are new skills. They were always the senior skills. What changed is that they are now the whole job rather than the last twenty percent of it. ## Verification is the new typing When I review generated code, I read it slower than I read human code, not faster. Human code has a consistent author with consistent blind spots. Generated code is plausible everywhere and wrong in places that do not look wrong: an off-by-one in a pagination loop, a retry without backoff, an error swallowed in a catch block that logs and continues. In systems I audit, these are the failures nobody else finds, because everyone assumed the code that compiles and passes the happy path is done. The engineer who can find them is worth more than before, not less. ## The junior problem I worry about how the next generation learns. If the model writes the first draft, where does a junior get the ten thousand small mistakes that build intuition? My answer, for people I mentor, is deliberate: write it yourself first, then ask the model, then diff. Read the model's version as a reviewer, not as a consumer. Keep a list of the classes of bugs it produces. The repetition is still available, it is just no longer forced on you, and you have to choose it. ## Where to invest - System design, because deciding boundaries and data flow is still a human job and the model amplifies whatever design it is given. - Testing and evals, because generated code needs more verification, not less. - Domain depth, because the model knows the general pattern and you know what is actually true in your business. - Writing, because a precise specification is now executable, and the person who writes it well ships faster than the person who only codes well. The boilerplate is gone and I am glad. What replaced it is harder and more interesting, and it rewards exactly the habits that made good engineers good before any of this existed. --- ## DNS TTLs and other small things that bite at 3 a.m. 2026-02-27 · infra The big architectural failures get the post-mortems and the conference talks. The failures that actually wake people up are smaller. A number in a configuration file. A default nobody read. A certificate with a date on it. In the audits I do, this category produces more incidents than any design flaw, and it is the easiest to prevent because every item on the list is known in advance. ## DNS TTL, the migration you cannot undo quickly A DNS record's time to live tells resolvers how long to cache the answer. A TTL of a day means that after you change the record, some users will keep hitting the old address for up to a day, and you cannot make them stop. This is fine on a normal Tuesday. It is a disaster during a migration or a failover, when the whole point is that traffic moves now. The discipline is simple and almost nobody does it. A week before any planned change, drop the TTL to a minute or five. Wait for the old TTL to expire everywhere. Make the change. Confirm. Raise the TTL back to something reasonable, an hour or so, because very short TTLs increase load on your DNS provider and add latency to every cold lookup. The TTL you need during a failover has to be set before the failover, and a failover is by definition unplanned. So the record that fronts your production origin should probably stay at five minutes permanently. ## Certificates, the calendar bomb Certificates expire on a date that was known when they were issued. There is no excuse for being surprised, and yet expired certificates take down a remarkable share of the systems I have seen. The usual cause is that automatic renewal was configured for the main domain and forgotten for the internal one, the API subdomain, the mail server, or the certificate pinned inside a mobile app. The fix is an inventory: every certificate in the system, where it is used, who renews it and how, and an alert thirty days before expiry that goes to a human, plus a second alert at seven days that goes to a louder human. Automate renewal wherever the platform supports it, and then verify the automation actually renewed, because automation that silently fails is the same as no automation. ## Timeouts that do not match A load balancer with a sixty-second idle timeout in front of an application with a ninety-second request timeout means that every request between sixty and ninety seconds returns an error to the user while the application keeps working on it. The application logs a success. The user sees a failure. The support ticket says 'it sometimes works'. Every hop in the request path has a timeout: browser, CDN, load balancer, application server, database driver, database. They must decrease as you go deeper, so that the innermost gives up first and the layers above can handle the failure cleanly. Write them down in a single table. When one changes, check the table. ## The rest of the list - Connection pool size smaller than the number of concurrent workers, which manifests as random slowness under load and nothing in the logs. - Disk full on a database volume because logs, temp files or WAL archives grew without a retention rule, which manifests as everything failing at once. - Clock drift on a server that was never configured to sync time, which manifests as token validation failures nobody can reproduce. - Rate limits on a third-party API that were fine at launch and are exceeded now, which manifests as intermittent failures around the top of the hour. - A cron job that runs in UTC on a server that was assumed to be in local time, which manifests as reports arriving three hours early. None of these are interesting. All of them are on my checklist for every system, because the interesting failures are rare and these are not. ## The habit Keep one document per system listing every number that could bite: TTLs, certificate dates, timeouts by layer, pool sizes, disk retention rules, rate limits, time zones. Review it quarterly. It takes an hour. It is the highest-yield hour in operations, and it is the difference between an incident at 3 a.m. and a calendar reminder at 3 p.m. --- ## State machines for everything that matters 2026-02-24 · system-design, engineering Every system I audit has an order, a subscription, a document, a job or a payment that moves through states. And most of them represent that lifecycle as a pile of boolean columns. is_paid, is_shipped, is_cancelled, is_refunded. I call this flag soup, and it is where a surprising share of production bugs live. The trouble with flags is that they can all be true at once. An order that is shipped and cancelled and refunded and not paid is not a state. It is a contradiction that the database happily stores and that some code, somewhere, has to interpret. Every report, every screen, every job then encodes its own private theory about what those combinations mean. ## Make the states explicit A state machine replaces the soup with two things: a finite list of states and a finite list of allowed transitions. An order is in exactly one of draft, pending_payment, paid, shipped, delivered, cancelled, refunded. The transition table says paid can go to shipped or refunded. It does not say cancelled can go to shipped. The application enforces that table. Any attempt to make an invalid transition is rejected with an error, not silently applied. That single rule kills an entire class of bugs, because the impossible states become impossible in fact, not just in intention. In practice this is a status column plus a small transition function that is the only code allowed to change it. No direct UPDATE orders SET status from random services. One door in. ## Record every transition The status column tells you where the order is. It does not tell you how it got there. So alongside it, keep a transitions table: order_id, from_state, to_state, timestamp, actor, reason. Append only. This is your audit trail, and it pays for itself the first time a customer asks why their order was cancelled. It also gives you free metrics. How long do orders sit in pending_payment? What fraction go from shipped to refunded? Those are just queries over the transitions table. With flag soup, the same questions require archaeology. ## Terminal states and timeouts Some states are terminal. Delivered, refunded, cancelled: nothing leaves them. Say so in the table. Code that tries to "reactivate" a refunded order should fail loudly. Timeouts should be transitions too. An order in pending_payment for more than 30 minutes moves to expired. Not by a flag flipping somewhere, but by a scheduled job that performs the same transition through the same door as everything else. That way the timeout is logged, audited and reported exactly like a human action. ## Why this makes bugs obvious When the lifecycle is explicit, bugs show up as invalid transition errors in your logs instead of as weird rows discovered months later. A report that groups by status is trustworthy, because status means one thing. A new engineer can read the transition table and understand the business in ten minutes. It also changes conversations with product. "Can a delivered order be cancelled?" becomes a question about a row in a table, not an investigation across five services. ## TypeScript fits this naturally In TypeScript, a discriminated union is a state machine waiting to happen. Each state is a variant with its own fields: a shipped order carries a tracking code, a refunded order carries a refund id and a reason. The compiler refuses to let you read the tracking code on a draft. Pair that with a transition function whose signature says which states it accepts, and exhaustive switch checks that fail compilation when you add a state and forget a branch. The type system becomes the enforcement layer at build time, the transition table at runtime, and the transitions log afterwards. Whenever something in your domain has a lifecycle, and almost everything that matters does, model it as a state machine. It costs one table and a little discipline. Flag soup costs you every week, forever. --- ## What I look for in an engineer 2026-02-22 · career I have reviewed the work of a great many engineers across 600+ projects, and I have hired for my own studio. Over time the things I look for have simplified. It is not the stack. It is not the years. It is a small set of habits that show up in the first hour of working with someone and predict most of what happens after. ## They read first, and admit what they do not know The first thing I watch is what an engineer does when handed an unfamiliar system. The weak signal is opening the editor and starting to type. The strong signal is opening the logs, the README, the git history, the schema, and asking who built this and why. Engineers who read first make fewer changes and better ones. They also find the failures nobody else finds, because most failures are visible in the history if you look. Then I ask questions I know are hard. I am not testing knowledge. I am testing what happens at the edge of knowledge. The engineer I want says "I do not know, let me check", checks, and comes back with an answer and a source. The one I avoid guesses confidently. On a production system, a confident guess is more dangerous than an honest gap, because the gap gets filled and the guess gets deployed. ## They notice the small wrong thing A field named user_id that holds an email. A timeout set in the config file but never read by the code. A test that passes because the assertion was commented out. I care about whether an engineer sees these, not because each one is important, but because the habit of seeing them is what stops the big failure. The big failure is always a small wrong thing that was visible for months. Detail-obsessed is a compliment I take seriously, and it is the trait I look for first. ## They finish Starting is common. Finishing is rare. Finished means deployed, documented, handed to whoever will maintain it, and observed in production long enough to know it works. I look at someone's history for things that were completed, especially unglamorous things. A person who finishes a boring migration will finish the interesting project. A person with ten exciting half-built repositories will have eleven next year. ## They can explain a decision to someone who does not code At some point every engineer must tell a founder, a product manager, or a client why the thing they want is a bad idea, or why the delay is real. The ones who can do that in plain language, without jargon and without condescension, become the people the business trusts. The ones who cannot end up isolated, correct and ignored. I ask candidates to explain a past technical decision to me as if I were a customer. The quality of that explanation tells me more than their code. ## What I do not weigh, and how to show what I do I care less than most people about the specific framework. I prefer Next.js and TypeScript for my own work, but an engineer with the habits above will learn the stack in a month. I care little about the degree; I did not finish mine. I care little about the size of the previous company. Some of the sharpest people I have worked with came from a shop floor or a repair bench, and they had the reading habit and the noticing habit before they ever touched a compiler. If you are being evaluated, here is how to show these things in an interview or a trial week. Read the system before you touch it and say what you found. Say "I do not know" out loud at least once and then come back with the answer. Point out one small wrong thing you noticed. Talk about something you finished, including the boring last mile. Explain one past decision in words your grandmother would follow. If you are the one evaluating, stop asking trivia. Hand them something real, watch the first hour, and listen to how they explain themselves. Everything else is noise. --- ## The deploy is part of the design 2026-02-20 · infra, system-design I have reviewed a lot of architecture documents. Most of them stop at the point where the code is written. Boxes, arrows, a database, maybe a queue. Then a single line: 'deployed via CI/CD'. That line hides half the risk in the system. The deploy is not a separate concern that DevOps handles later. It is a design decision with the same weight as choosing the database. A system that cannot be released safely is a system that will be released rarely, and a system released rarely accumulates large, frightening changes. The design determined that outcome long before anyone touched a pipeline. ## Questions the design must answer When I review a design now, I ask deployment questions in the same breath as data questions. Can two versions of this service run at the same time? If not, every deploy is downtime, and the design has decided that for you. Does the schema change ship before or after the code that uses it? If the answer is 'at the same time', the design has a hidden ordering problem. What does a half-finished deploy look like? Half the instances on the new version, half on the old, sharing a cache with a new key format. Did anyone draw that state? These questions are not operational trivia. They shape the code. A service designed to tolerate two concurrent versions writes its events with a version field, reads unknown fields tolerantly, and never renames a column in the same release that stops writing to it. Those are design choices, and they are cheap on day one and expensive on day four hundred. ## The stateful parts decide everything Stateless services are easy to deploy. Roll them, they come back. The hard parts of every deploy are the stateful ones: the database, the cache, the queue consumers, the long-running jobs, the WebSocket connections. Each of those needs a plan that is part of the design. For the database, the plan is expand-then-contract: add the new column, backfill, switch reads, switch writes, drop the old column, each step its own release. For queue consumers, the plan is that new messages must be readable by old consumers for the duration of the rollout, which means additive schemas only. For long jobs, the plan is that a job must be resumable, because the deploy will kill it in the middle. For WebSockets, the plan is that clients reconnect with backoff and the server drains connections before it exits. If the design does not include those plans, the team will invent them under pressure, and the invented version usually involves a maintenance window at midnight. ## Rollback is a feature you build 'We can just roll back' is the most common false statement in deployment planning. Rolling back code is easy. Rolling back a migration that dropped a column is impossible. Rolling back a service that already published events in a new format is a data repair project. So I design for rollback explicitly. Every change is either backward compatible or it is split into compatible steps. The question 'if we revert this in an hour, what breaks?' is asked in the review, not in the incident. The honest answer is sometimes 'this step cannot be reverted', and that is fine as long as it is known, tested twice, and shipped alone on a quiet morning. ## Make the pipeline enforce the design Once the deploy strategy is part of the design, the pipeline becomes its enforcement. Migrations run in a separate step that must succeed before the application deploys. Health checks verify the new version can talk to the database before traffic arrives. A smoke test hits the three most important endpoints. A canary receives a small share of traffic for a fixed period before the rest follows. None of this is exotic. All of it is cheap when it is part of the original shape of the system. The teams I see struggling are the ones who designed a beautiful system and then tried to deploy it as an afterthought. The deploy was always part of the design. They just did not draw it. --- ## The bug is always in the boundary 2026-02-19 · audit, architecture Inside a well-written function, bugs are rare. The author held the whole thing in their head, the types line up, the tests cover it. The serious bugs, the ones that cost money and reach the incident channel, live somewhere else: at the seams. Where one service calls another. Where a string becomes a number. Where the browser's clock meets the server's. Where the code meets the human. After hundreds of audits I no longer look for bugs in the middle of things. I look for the edges and read them twice. ## What a boundary is A boundary is any place where two parts of a system hold different assumptions about the same data. A service that stores amounts in cents calls one that expects decimals. A frontend that sends dates as local time talks to a backend that stores them as UTC and forgets to convert. A queue delivers messages at least once to a consumer written as if exactly once. A form field trims whitespace and the database unique index does not. Each side is correct on its own terms. The bug is the disagreement, and nobody owns the disagreement because it is not inside anyone's code. ## The boundaries I check first Types crossing the wire: every JSON payload is stringly typed until proven otherwise, and 'true', true and 1 are three different values that get treated as one. Time: any place a timestamp is created, parsed or compared, with special attention to whether the timezone is explicit. Money: any place a currency value changes representation, especially float to anything. Identity: any place an ID from one system is used to look something up in another. Delivery semantics: any place a message, webhook or event is consumed, and whether the consumer survives receiving it twice or out of order. Authorization: any place the code trusts that the caller already checked who the user is. In a marketplace I audited, the entire pricing bug came down to one boundary: a total computed as a float in one service and stored as an integer in another, with the rounding happening on different sides depending on the code path. Every module was correct. The seam was not. ## Why boundaries are under-tested Unit tests live inside a module, by definition. Integration tests exist but are slower, flakier and fewer, so they cover the happy path across the seam and nothing else. Nobody writes the test for what happens if the other side sends a null here, because the other side's contract says it never will, and the contract is a comment. The result is that the parts of the system with the most assumptions have the least verification. ## How to make boundaries safe Validate at every boundary, on entry, with a schema, and reject what does not conform. Do not trust the other side even when it is your own code, because your own code changes. Convert to a canonical internal representation immediately: one type for money, one for time, one for identifiers, and never let the wire format leak inward. Write the contract as code, not prose, and version it. And test the boundary explicitly: the duplicate message, the missing field, the timezone offset, the string where a number was expected. These tests are cheap to write once you decide the boundary is where the risk is. ## The human boundary The last seam is the one between the system and the person using it. The field that accepts a date in two formats and silently picks the wrong one. The button that can be pressed twice. The confirmation that looks like a success when it was a partial failure. Every one of those is a boundary with the same properties as the ones between services, and the same rule applies: assume the other side will do the unexpected thing, and make sure the system stays correct when it does. --- ## TypeScript as a design tool, not a linter 2026-02-17 · fullstack, typescript Most teams use TypeScript as a spell checker. It catches a misspelled property, a missing argument, an undefined that slipped through. That is useful, and it is a tenth of what the tool can do. I use TypeScript the way I use a whiteboard. Before I write a function body, I write the types. Before I decide how a feature behaves, I decide what states it can be in and I make the compiler refuse every state I did not list. ## Make the impossible unrepresentable Here is the difference in practice. A naive order type has a status string, an optional paidAt date, an optional shippedAt date, an optional cancellationReason. Every combination is legal to the compiler, including a cancelled order that was shipped and never paid. The designed version is a discriminated union. A pending order has no dates. A paid order has paidAt and nothing else. A shipped order has both dates. A cancelled order has a reason and a timestamp. Now a function that renders the tracking page cannot even be written for a cancelled order, because the shippedAt field does not exist on that branch. That is not a linting improvement. That is a design decision, encoded in a place where it cannot drift from the code. The next engineer who adds a refunded status gets a compiler error in every switch that forgot about it, if you close those switches with an exhaustive `never` check. ## Types as the first draft of the contract When I start a feature, I write the type of the input, the type of the output and the union of everything that can go wrong. A function that returns a result type with an explicit error union tells the caller more than any docstring. It tells them that a payment can be declined, that the card can be expired, that the provider can time out, and it forces them to handle each one or explicitly ignore it. I do this before I know how the function will be implemented. Often the types reveal that the feature is more complicated than the ticket suggested, and that is the moment to have that conversation, not after two weeks of code. Branded types are the next tool. A user id and an order id are both strings at runtime, and I have found real bugs where they were swapped. A brand is a zero-cost type-level tag that makes the swap a compile error. One line of type definition, and an entire category of mistakes is gone. ## Where the compiler earns its keep The strict flags are not optional in my projects. `noUncheckedIndexedAccess` alone has caught more bugs in audits than any test suite I have read, because it forces you to admit that an array lookup can miss. The `satisfies` operator lets me check that a configuration object matches a schema without widening its type, which keeps autocomplete precise. Template literal types let me type route strings and event names so that a typo in an event name is a compile error rather than a silent no-op in production. I keep types close to the data they describe, and I derive types from the source of truth instead of duplicating them. If the database schema is the truth, generate the types from it. If a validation schema is the truth, infer the type from the schema. Two definitions of the same shape will disagree eventually. ## The cost, and when to stop Type-level design has a ceiling. When a type takes longer to read than the function it protects, you have crossed it. Deeply recursive conditional types slow the compiler, confuse the team and rarely prevent a bug that a simpler type would not. My rule is that a type should be readable by an engineer who has been on the team for a week. If it needs a comment explaining the trick, it is a trick, and tricks do not survive turnover. But inside that limit, the compiler is the cheapest reviewer you will ever hire. It reads every line, it never gets tired, and it does not care that the deadline is Friday. Design with it, not around it. --- ## How I review code a model wrote 2026-02-13 · ai, audit A model-written pull request looks better than most human ones. Consistent naming, tidy structure, comments in the right places, a test file that exists. That polish is the first problem, because reviewers relax when the code looks competent. I do not review generated code more leniently and I do not review it more harshly. I review it with a different map of where the bugs are likely to be, because the distribution has changed. ## Same standards, different distribution of bugs Human bugs cluster around fatigue and haste: the off-by-one at the end of a long day, the copy-paste that was not updated. Model bugs cluster around assumptions: the code is a confident implementation of a task as the model understood it, and the misunderstanding is usually at the boundary between the task and the rest of the system. The function is correct. The call site, the error path, the concurrency story and the data contract are where it guessed. So that is where I spend the review. ## What I read first I read the error handling before the happy path. Generated code loves to catch an exception, log it and continue, which turns a loud failure into a silent one. I read every place a value crosses a boundary: an API response parsed without validation, a database row assumed to have a field, a date parsed without a timezone. I read anything that touches money, identity or deletion twice. And I read the imports, because models will reach for a library that is plausible, deprecated, or not the one the project already uses for the same job. ## The patterns that recur - Swallowed errors: a try-catch that hides a failure the caller needed to know about. - Invented interfaces: a call to a method that does not exist on that object, or exists with a different signature. - Missing idempotency: a handler that is fine once and wrong on retry. - Optimistic parsing: trusting external input because the example in the prompt was well-formed. - Duplicated logic: a new helper that reimplements something the codebase already had, slightly differently. ## Tests written by the same model Generated tests deserve special suspicion because they were produced by the same process that produced the code, from the same understanding, with the same blind spots. A test that asserts the code does what the code does is tautological. I look for whether the tests include the case the author of the task would care about: the empty input, the duplicate, the timeout, the unauthorized caller. If those are absent, the tests are documentation, not verification, and I write the missing ones myself before approving, or ask the model to, with the specific case named. ## The reviewer is the owner Whoever approves the merge owns the behavior. That was always true and it is more true now, because the author of the code cannot be asked why they made a choice, and if it can, it will make up a reason. So the review is the moment the code acquires an owner. I do not approve anything I could not explain at an incident review. If a generated change is too large to hold in my head, I ask for it in pieces, exactly as I would from a person. The model's speed is real. The review has to be as slow as it always was, because the thing being reviewed is not the code. It is the understanding of the person about to sign it. --- ## Eventual consistency, explained to the product team 2026-02-11 · system-design A product manager once asked me why a user could create an order, open the orders list, and not see it for three seconds. She was not asking for a lecture on distributed systems. She wanted to know if it was a bug. The honest answer is: it depends on what we promised, and we never wrote the promise down. This article is the version of that conversation I wish I had given the first time. ## What "eventual" really means In a system with more than one copy of the data, writes go to one place and reads may come from another. The copies catch up, usually in milliseconds, sometimes in seconds, occasionally in minutes when something is unhealthy. "Eventually consistent" means the copies will agree, but not right now. The orders list showing the order three seconds later is not the database lying. It is a read served from a replica, a search index, or a cache that has not yet received the write. Nothing was lost. The user just looked at a copy that was behind. ## The two guarantees users actually feel Users do not care about consistency models. They care about two specific experiences, and both have names. Read-your-writes: if I just did something, I should see it. I created the order, so my orders list must include it. Breaking this makes users think the action failed, so they do it again, and now they have two orders. This is the guarantee that matters most for anything the user just touched. Monotonic reads: I should never see time go backwards. If the list showed five items and I refresh, it should not show four. Breaking this happens when consecutive reads hit different replicas at different lag. It looks like data disappearing, and it destroys trust faster than a slow page. Both can be provided without making the entire system strongly consistent. Route the user's own reads to the primary for a short window after a write. Pin a session to one replica. Or, simplest of all, update the UI from the response you already have instead of re-fetching. ## Which flows can lag and which cannot Not every flow deserves the same guarantee, and pretending it does is how systems get slow and expensive. The rule I use: anything that moves money, reserves a scarce resource, or grants access must be strongly consistent at the moment of the decision. Charging a card, reserving the last seat, decrementing inventory, checking a permission. These read from the source of truth, under a lock or a transaction, and they accept the latency cost. Almost everything else can lag. Analytics dashboards, the search index, the notification that says "your order shipped", the recommendation feed, the count of unread messages. If a search result is two seconds behind, nobody is harmed. If inventory is two seconds behind, you sell the same item twice. When a product team classifies flows this way, engineering gets to spend consistency where it matters and speed everywhere else. ## Designing a UI that tells the truth The worst option is a UI that pretends. It shows a spinner, the request returns, the list re-fetches from a lagging replica, and the item is missing. The user is now confused and the system did nothing wrong. Optimistic UI is the better default for user-initiated actions: show the item immediately from the data you already have, mark it subtly as pending if the backend has not confirmed, and reconcile when the confirmation arrives. If the write fails, remove it and say so. Pending states are honest. "Payment processing" is better than a green checkmark that gets revoked. "Your report will be ready in a moment" is better than an empty dashboard that fills itself in silently. The promise to write down is simple: for each screen, what the user just did must be visible immediately, what others did may be a few seconds behind, and nothing ever goes backwards. Ship that, and the three second question stops being a bug report. --- ## Cardiovascular data is a systems problem 2026-02-08 · science, data An electrocardiogram is a waveform sampled hundreds of times per second. An echocardiogram is a video with measurements a human typed next to it. A lipid panel is a handful of numbers from a lab, on a date that may or may not match the visit. A blood pressure reading is one pair of numbers, taken by a nurse or a wrist device, under conditions nobody recorded. A clinical note is free text. A wearable produces a heart rate every few seconds for months. Each of these is cardiovascular data. None of them agree on what a row is. ## Time is the first boundary Every one of those sources has its own clock. The ECG has milliseconds. The lab has a collection date and a result date, which differ. The note has the time the doctor wrote it, not the time the thing happened. The wearable has the phone's timezone, which changed when the patient traveled. Before you can ask a single scientific question that spans two sources, you have to decide what 'at the same time' means, and that decision is an engineering decision with scientific consequences. Align on the wrong clock and you will find correlations that are artifacts of the alignment. I have audited enough systems with timezone bugs to know this is not hypothetical. It is the most common way two correct datasets combine into one wrong one. ## Identity is the second Who is this patient? The hospital has one identifier, the lab has another, the device has a serial number tied to an account, the study has an anonymized code. Linking them is required for the science and restricted by ethics and law, which means the linkage itself is a system: keyed, audited, access-controlled, reversible only by the people allowed to reverse it. Get identity wrong in one direction and you merge two people. Get it wrong in the other and you split one person into two, which quietly halves your sample and doubles your noise. ## Missingness is not empty A missing value in a cardiovascular dataset is rarely random. The troponin was not measured because nobody suspected an event. The echo was not done because the patient was too unstable to move. The wearable stopped recording because it was charging, or because the person was in hospital. Every gap carries information about the process that produced it, and any analysis that treats gaps as neutral is measuring the process, not the heart. An engineer's job here is not to impute. It is to preserve the reason for the gap as data, so the scientist can decide what it means. ## Units, versions and the silent reformat A device firmware update changes how it filters noise. A lab switches assays and the reference range shifts. A hospital migrates its records system and a field that used to be mmHg is now stored as a string with the unit inside it. None of this is announced to the researcher downstream. In software we would call these breaking changes to an API and demand a version bump. In clinical data they happen silently, and the only defense is to record the version of everything, every time, and to check distributions when they drift. ## What this means for building tools When people hear 'AI for cardiovascular science' they picture the model. I picture the long months of plumbing before a model can be trusted with anything: time alignment, identity resolution, missingness tracking, unit normalization, versioning, provenance on every derived value. That plumbing is where the errors hide, and it is where I have spent most of my career finding them. It is also why ELUCENIA is built with clinicians and scientists, not for them. They know what the data means. I know how it breaks in transit. Neither is enough alone, and the field deserves both. --- ## The edge runtime: when it helps and when it just moves the problem 2026-02-05 · fullstack, nextjs The pitch for edge compute is simple: run your code in the data center closest to the user and every response gets faster. The pitch is true for a narrow class of code and misleading for most of what a product does. I have watched teams move an application to the edge, celebrate the demo and then quietly move it back three months later. The reason is geography. Your code moved. Your database did not. ## What the edge runtime actually is In Next.js, the edge runtime is a restricted JavaScript environment built on Web APIs. No Node.js modules, no file system, no raw TCP sockets, a small bundle limit and very fast cold starts. It runs on a CDN's points of presence around the world instead of in one region. Those restrictions are the point. Because the runtime is small, it starts in milliseconds anywhere. Because it cannot open a socket, it cannot use a traditional database driver, and it talks to data over HTTP. That is fine for a key-value lookup and a serious constraint for anything relational. The framework itself has become more careful about this. The request interception layer that used to be edge-only now runs on the Node.js runtime by default, precisely because most teams needed database access there and the edge could not give it to them. ## When it helps The edge shines when the decision can be made with what is already in the request. Redirecting by country. Rewriting a URL based on a cookie for an A/B test. Verifying a signed session token, without a database lookup, to bounce anonymous users before they reach the origin. Adding security headers. Serving a personalized variant of a cached page where the personalization is a header value. In every one of those, the code needs the request, maybe a small globally replicated config store, and nothing else. The response is faster because the round trip to the origin region is gone, and that round trip is 100 to 300 milliseconds for a user on the other side of the planet. It also helps for static content with light dynamic touches, such as a landing site that swaps currency by region, where the heavy content comes from the CDN cache and the edge function only decorates it. ## When it moves the problem Now take a product page that runs three queries against a database in a single region. On the origin server, those queries are each a millisecond of network time because the database is in the same building. On the edge, each query crosses the ocean. Three sequential queries from an edge location far from the database cost more than the entire origin render would have. You have not made the page faster. You have moved the latency from the user-to-server hop to the server-to-database hop, multiplied it by the number of queries, and made it harder to debug. This is the failure I see most often, and the demo hides it because the demo runs from the same region as the database. Globally replicated databases solve this on paper and create a consistency problem in practice. A write in one region takes time to reach the others, and an application that reads its own write from a different region sees the old value. That is a real trade-off some products should make. It is not something you should discover after moving to the edge. ## The decision rule I use one question. Does this code need to talk to a single-region data store more than once per request? If yes, run it in the origin region, next to the data, and use the CDN for what it is good at. If no, the edge is a good home for it. Most of an application answers yes. The pieces that answer no are small, valuable and worth putting at the edge on purpose. The mistake is moving the whole thing and hoping the geography sorts itself out. It does not. Data has a location, and the compute should be near it. --- ## Reading legacy code like an archaeologist 2026-02-03 · architecture, audit I have read more code written by strangers than code I wrote myself. After a few hundred systems, I stopped reading legacy code as an engineer looking for what is wrong and started reading it as an archaeologist looking for what happened. The second approach finds more bugs, and it finds them faster. An archaeologist does not judge a ruin for not being a modern building. They ask what it was for, who built it, what changed and why the layers are in the order they are. Legacy code deserves the same respect, not because it is good, but because the story it tells is the fastest route to understanding what will break when you touch it. ## Stratigraphy: the layers tell the story Every codebase older than three years has visible strata. The original style, usually consistent and naive, from when one or two people wrote everything. A second layer with different naming and a new abstraction, from when the team grew or a new lead arrived. A third layer of patches that ignore both previous styles, from the period when nobody had time. And sometimes a fourth, a half-finished migration to something modern, abandoned mid-way. I map these before I read a single function in depth. Git blame by directory, sorted by date, does most of the work. The transition points between layers are where the surprises live. That is where two conventions meet, where the old error handling hands off to the new, where a data shape got translated and something was lost. ## Artifacts: the strange things are the important things The archaeologist's most useful instinct is that anything weird was probably put there on purpose. A hardcoded sleep of 350 milliseconds. A retry loop that tries exactly four times. A check for a customer id that appears nowhere else. A commented-out block with a date from six years ago. Junior engineers delete these as cleanup. I treat each one as evidence of an incident that nobody documented. The sleep exists because some downstream system needed a moment, and it still might. The four retries were calibrated against something. The special customer id was a production hotfix for someone important. Before removing any of them, I try to find the incident. Commit message, ticket reference, a comment in a nearby file. If I cannot find it, I do not remove it; I wrap it in a comment saying it is unexplained and add monitoring around it. ## Provenance: who wrote it and under what pressure Code written in the week before a launch looks different from code written in a quiet month, and the difference matters when you estimate how safe it is to change. I look at commit density around each module. A burst of thirty commits in three days followed by silence for two years is a feature that was rushed and then never touched. It works, in the sense that nobody has complained, but it has probably never been exercised outside its happy path. The reverse pattern, steady small commits over years, is a module that has been maintained, understood and probably tested by the people who kept it alive. Those are safer to change, even if the code is uglier, because the ugliness has been reviewed many times. ## Reconstruction: rebuild the intent before rebuilding the code The final step, before proposing any change, is to write down what the code is trying to do in plain sentences, as if explaining it to the person who wrote it. Not what it does line by line; what it is for. 'This job reconciles yesterday's payments against the provider's report and flags mismatches over a threshold.' If I cannot write that sentence, I do not understand the module well enough to change it. Then I list the questions the code raises that I could not answer from the code alone. Why is the threshold a percentage here and a fixed amount there? Why does this path skip the audit log? Those questions go to whoever has been there longest, and the answers usually reshape the plan. Legacy code is the most honest documentation a company has. It records every decision that actually shipped, including the ones nobody would admit to in a design review. Read it like a record, not like a mess, and it will tell you exactly where it is going to fail next. --- ## Agents need boundaries too 2026-01-31 · ai, architecture Strip away the vocabulary and an agent is a loop: observe, decide, call a tool, repeat until done. That is not new. Cron jobs, workflow engines and retry loops have been doing it for decades, and the discipline we built for them, boundaries, permissions, timeouts, idempotency, still applies. What is new is that the decision step is a language model, which means the loop can take actions nobody wrote down in advance. That makes boundaries more important, not less. ## An agent is a loop with permissions When I review an agent, the first thing I draw is not the prompt. It is the list of tools, and next to each tool, what it can touch. Read a document. Search an index. Send an email. Write to the database. Call a payment API. That list is the agent's real capability, regardless of what the prompt says it should do. If a tool can delete records, the agent can delete records, and a sufficiently strange input will eventually make it try. Prompts are guidance. Tools are permissions. I design the second and hope for the first. ## The blast radius of a tool The rule I use: every tool gets the narrowest permission that lets the agent do its job, and every tool with a side effect is wrapped in something the agent cannot bypass. A support agent that can issue refunds gets a refund tool with a cap, a per-customer rate limit and a log, not access to the billing service. A coding agent gets a sandboxed working tree, not my shell. A research agent gets read access to a corpus, not the internet. When the agent misbehaves, and it will, the damage is bounded by the tool, not by the model's good intentions. ## Boundaries I put around every agent - A budget: maximum steps, maximum tokens, maximum wall-clock time. The loop terminates by construction, not by hope. - Idempotent tools with keys, so a retried step does not repeat a side effect. - A human gate on any irreversible action, with the proposed action rendered for a person to approve. - Structured tool inputs validated in code before execution, because the model will eventually produce a malformed call. - A trace of every step, stored, so I can replay what it decided and why. ## Stopping is a feature The most dangerous agent I have reviewed was one that was very good at not giving up. It would retry, rephrase, try another tool, try the first tool again with slightly different arguments, for as long as it was allowed. On a read-only task that is charming. On a task with side effects it is a machine for producing duplicates. An agent needs a clear notion of done, a clear notion of stuck, and a path that ends in asking a person rather than trying harder. I would rather have an agent that stops early and reports than one that finishes at any cost. ## Same rules, older names None of this is exotic. Least privilege, bounded retries, compensation for side effects, audit logs, human approval for irreversible actions. We learned these building payment systems and deployment pipelines, and we learned them by getting burned. Agents just make the lessons urgent again, because the thing inside the loop is creative and the things around the loop had better not be. Put the creativity in the middle and the boundaries on the edges, and you can let the model be as clever as it likes. --- ## The API route that became a monolith 2026-01-29 · fullstack, nextjs, architecture I can describe the file before I open it. It is called route.ts, it is between 600 and 1,200 lines long, and it started as a fifteen-line handler that saved a form. Somewhere along the way it acquired a switch on the request body, three database clients, an email call, a webhook to a CRM, and a comment that says 'temporary'. I have read that file in dozens of codebases. It is the most common architectural failure in Next.js applications, and it happens because the framework makes it easy to put code somewhere without deciding where it belongs. ## How it happens The App Router gives you route handlers and Server Actions as places to run server code. Both are entry points. Neither is a place for business logic, but nothing stops you from putting it there, and the first version is always small enough that it seems fine. Then the second feature needs the same validation, so it gets copied. The third feature needs a slightly different version, so the handler gets a parameter. The fourth needs to run after the third, so they get chained with a flag. Each step is a reasonable local decision. The sum is a monolith with no boundaries, living inside a file whose only job was to parse a request. The tell is when the route handler is imported by something else. An entry point should have no callers except the framework. The moment another module needs the logic inside it, the logic is in the wrong place. ## The structure that holds My rule is that an entry point does exactly four things: authenticate, validate input, call one function, and shape the response. Ten to thirty lines. If it is longer, something has leaked in. The function it calls lives in a domain module: a plain TypeScript function that takes typed input and returns a typed result, and knows nothing about HTTP, forms or the framework. It can be called from a route handler, from a Server Action, from a queue worker and from a test, and it behaves the same way in all four. That separation is the whole trick. The route handler and the Server Action for the same operation become thin wrappers around the same domain function. When a mobile app needs an API, you add a third wrapper. When a cron job needs to run the operation nightly, a fourth. The logic never moves. Below the domain module sits the data access layer. I keep queries in their own modules, named for the aggregate they touch, and the domain function calls them by name. The domain never builds a query. That is what keeps the ORM from spreading into every file. ## Where side effects go The monolith route usually contains side effects inline: send the email, call the CRM, update the search index. Each one adds latency to the request and a new way for it to fail halfway through. Side effects belong in an outbox or a queue. The domain function writes the record and an event row in the same transaction. A worker reads events and performs the effects, with retries. The request returns as soon as the record is committed. The user does not wait for the CRM, and the CRM being down does not fail the checkout. In a small project, the worker can be a route handler triggered by a scheduler. In a bigger one, it is a real queue consumer. The domain function does not change either way. ## Refactoring the one you already have Do not rewrite it. Extract from the inside out. Find the smallest self-contained block, usually validation, and move it to a module with a test. Then the next block. The route gets shorter with every pull request, and it keeps working the entire time. I set a lint rule that fails when a file under the app directory exceeds a line count, and I make it strict enough to hurt. Limits that hurt get respected. Limits that are comfortable get ignored, and the file grows back. The monolith is not a Next.js problem. It is a placement problem. Decide where the logic lives before you write it, and the entry point stays what it should be: a door, not a house. --- ## Caching is a consistency decision in disguise 2026-01-27 · system-design, architecture Every time someone says "let's just add a cache" I hear a different sentence: "let's keep a second copy of this data and hope it agrees with the first one." A cache is not a performance feature. It is a consistency decision that happens to make things faster. The moment you have two copies of a value, you have to answer a question you were avoiding: when they disagree, which one wins, and for how long is the wrong one allowed to be served? ## The question I ask before any cache What is the worst stale value a user can see, and for how long? That single question decides the strategy. A product description that is ten minutes old costs nothing. An account balance that is ten minutes old can trigger a chargeback. A permission check that is ten minutes old is a security incident. Write the answer down. If the team cannot agree on how stale is acceptable, you are not ready to cache that data. You are ready to argue about it, which is cheaper than arguing in production. ## TTL, invalidation, write-through A TTL is the honest option. You are saying "this value may be wrong for up to N seconds, and I accept that." It is simple, it self-heals, and its failure mode is bounded. Most read-heavy data with low correctness cost belongs here. Invalidation is the ambitious option. You promise to delete the cached entry every time the source changes. The promise is easy to make and hard to keep, because writes happen in places you forgot about: batch jobs, admin panels, a migration script, a second service that shares the table. Every missed path is a stale value with no expiry. Write-through updates the cache and the store together. It gives you the freshest reads, but it turns the cache into part of your write path, so a cache outage now becomes a write outage unless you design around it. It is the right tool for small, hot, correctness-sensitive data, and the wrong tool for everything else. Many real systems mix them: invalidate on the paths you control, and keep a short TTL as a safety net for the paths you missed. That combination is not a cop-out. It is an admission that invalidation is never complete. ## Stampedes, negatives, and keys When a popular key expires, every request that arrives in the next few milliseconds misses at once and hits the database together. That is a cache stampede, and it usually shows up as a database spike exactly when traffic peaks. The fix is request coalescing: one in-flight fetch per key, and everyone else waits for that result. Stale-while-revalidate is the same idea with a nicer user experience, serving the old value while one worker refreshes it. Cache the misses too. If a lookup for a nonexistent user is expensive, and bots keep asking for nonexistent users, negative caching with a short TTL saves your database from answering the same "no" a thousand times. And be careful with what goes into the key. Tenant, locale, currency, feature flags, the user's role. If any of those affect the response and are missing from the key, you will eventually serve one tenant's data to another. I have seen this exact bug in more than one system I audited, and it always started with a key that looked complete on the day it was written. ## Treat the cache as a contract Write the staleness budget next to the code. Put the key format in one place. Add a metric for hit rate and another for the age of the served value. When someone asks why the page shows an old price, you should be able to answer with a number, not a shrug. A cache you cannot explain is a bug you have not found yet. --- ## Validation is not a phase 2026-01-24 · science, engineering Every mature software team has learned that quality is not a phase. You do not build for six months and then test. You test every commit, every merge, every deploy, and the tests are part of the product. Scientific software, and increasingly AI for science, still treats validation as something that happens at the end, in a paper, once. I think that is the single most dangerous habit in the space, and it is fixable with ideas we already have. ## What 'at the end' costs When validation is a phase, the cost of a wrong assumption grows with every step built on top of it. A cohort filter that silently dropped a subgroup gets discovered after the model was trained, the figures drawn and the manuscript written. Now the fix means redoing everything, and the incentive to look the other way is enormous. In software we know this price well: the bug found in production costs orders of magnitude more than the one found at the keyboard. ## Validation as a property of every artifact The alternative is that every artifact in the pipeline carries its own validation status. The raw data has integrity checks. The cleaned dataset has assertions about row counts, ranges and distributions that run when it is produced. The feature table has tests that compare it against a known sample. The model has an evaluation that runs on every retrain, against a held-out set that nobody trained on. The hypothesis the system surfaces has an evidence trail and a status: candidate, checked against literature, checked against data, validated by a scientist, replicated externally. Nothing moves to the next stage without its status. The status is data, stored with the artifact, visible in the interface. This is exactly how a CI pipeline works, applied to knowledge instead of code. ## The status must be honest and visible The failure mode I fear most in AI for science is not the wrong answer. It is the unvalidated answer presented with the same confidence as the validated one. Interfaces flatten everything into text, and text has no validation state. So the design rule is that status is never hidden: a candidate looks different from a validated result, in every view, every export, every place it can travel. If someone copies a number out of the system, the status travels with it. ## What validation means when the field is medicine In cardiovascular research the eventual consumer of a validated advance is a clinical decision. That raises the bar on what 'validated' means: not just that the code ran and the statistics were significant, but that the claim survived scrutiny by people who understand the biology and the clinical context, and ideally that it was reproduced by someone with no stake in it. That is slow, and it is supposed to be slow. My job is to make everything before that step fast and everything at that step unavoidable. ## The principle This is why 'every advance needs scientific validation' is one of the principles I hold ELUCENIA to, and why I mean it as an architectural constraint rather than a value statement. Validation is a field on every object, a gate on every transition, a visible mark on every screen. It is not a phase. A phase can be skipped when the deadline is close. A property cannot. --- ## Load testing the right thing 2026-01-22 · infra, performance A team shows me their load test results. Ten thousand requests per second, p99 under fifty milliseconds, the graph is flat and beautiful. Then I ask what endpoint they hit, and the answer is the homepage, or the health check, or a list endpoint with no filters that the CDN would have served anyway. The test measured how fast the framework can return a string. Production will fail on the checkout, the search, the report export, the endpoint that holds a transaction open while it calls a payment provider. Load testing is only as useful as the choice of what to load. Here is how I choose. ## Find the path that actually fails Start from the business, not the server. Which user action, if it failed under load, would cost the most? For a marketplace it is the checkout. For a clinic it is booking an appointment. For an analytics product it is the dashboard with a date range. That action is the test scenario, and the scenario includes every step a real user takes to get there: log in, load the page, add the item, apply the coupon, pay. The steps before the critical one matter because they warm caches, hold sessions and open connections in exactly the way production does. Then look at what that path touches. Database writes, not just reads. Locks on shared rows like inventory counts or account balances. Calls to external services with their own latency. Background jobs that get enqueued. A load test that skips any of those is testing a different system. ## Model the traffic, not the number 'Ten thousand requests per second' is not a traffic model. Real traffic has a shape: a mix of endpoints in known proportions, a ramp rather than a step, a small set of hot records that most users touch, and a long tail of cold ones. A test that hits one endpoint with random ids exercises none of the contention that real traffic produces. Build the mix from production logs. Take a busy hour, count requests by route, and replay those proportions. Take the distribution of record ids and preserve it, so the same popular products or accounts get hammered the way they are in reality. Ramp the load over minutes so you can see where the curve bends, rather than jumping to the target and seeing only that it broke. ## Test against production-shaped data A load test on a database with a thousand rows finds nothing. The query plans are different, the indexes fit in memory, the locks never contend. The test database must be the size of production, or a restored copy of it with sensitive fields scrubbed. This is the single most common reason load tests pass and production fails. The same applies to the caches. A cold cache and a warm cache are two different systems. Test both: the steady state with warm caches, and the moment after a deploy or a cache flush when every request goes to the origin at once. ## What to watch while it runs - Latency percentiles per endpoint, not overall, so a slow checkout is not hidden by a fast health check. - Error rate per endpoint, including the timeouts that the client counts as errors and the server logs as successes. - Database metrics: connections in use, lock waits, slowest queries, replication lag. - Saturation on the origin: CPU, memory, and especially the connection pool and the event loop delay. - Queue depth and job latency for anything enqueued by the tested path. The result you want is not a number. It is a sentence: 'the checkout holds p99 under 800 milliseconds up to 300 concurrent users, and beyond that the inventory row lock becomes the bottleneck'. That sentence tells you what to fix, at what traffic, and how much headroom you have. A flat graph of the homepage tells you nothing at all. ## Make it repeatable The first load test is a project. The tenth should be a command. Script the scenario, version it with the code, run it against a dedicated environment with production-shaped data, and keep the results next to each other so a regression is visible as a change in the sentence above. That is the difference between a load test as a one-time reassurance and a load test as an instrument. --- ## Contracts before code 2026-01-20 · architecture, engineering The most productive hour on any project I have led is the one where nobody writes implementation code. It is the hour where we write down what a component will accept, what it will return, what it will emit and how it will fail, and then argue about it before a single function body exists. This is not about documentation and it is not about waterfall. It is about doing the design where changes are cheap, in a text file, instead of where they are expensive, in three services and a migration. ## What a contract is A contract is the complete external behavior of a component, written so that two people who never speak could build the two sides. For an HTTP API that is an OpenAPI document. For an event, it is the schema of the payload with every field's meaning. For a module inside a monolith, it is the TypeScript interface of the public surface plus a paragraph per function saying what happens on failure. The parts that matter most are the ones teams skip. Which errors can this return and what does each mean to the caller? Which fields are optional now and which will always be present? What is the idempotency behavior when the same request arrives twice? What are the limits: page size, payload size, rate? What is the ordering guarantee, if any? A contract without those is a happy path with a name. ## Why before code Three reasons, and I have watched all three save weeks. First, the contract exposes disagreements while they are cheap. The frontend expected a list; the backend planned a paginated cursor. The mobile team needed the total count; the backend did not plan to compute it. Discovered in a contract review, that is a ten-minute conversation. Discovered after both sides shipped, it is a version bump and an apology. Second, the contract unblocks parallel work. Once the shape is agreed, the frontend can build against a mock, the backend can build against a test, and the integration test can be written by whoever is free. Without the contract, everyone waits for the slowest side. Third, the contract is the test. A schema can be validated at runtime. An OpenAPI document can generate a client and a set of contract tests that fail the build when the implementation drifts. The design artifact becomes the thing that keeps the implementation honest, permanently. ## The contract review The review is where the value is, and it should take less than an hour for most components. I ask a fixed set of questions. - What happens when each required input is missing, empty or malformed? - What does the caller see when the dependency behind this is down? - What is the response to the same request twice? - Which fields will be added later, and how will old callers survive them? - What is the largest payload, the longest list, the slowest case, and what happens at that limit? - Who consumes this, and did they read it? The last question is the one most often skipped. A contract nobody on the consuming side has read is a guess. I ask for one named consumer to sign off, even informally, before implementation starts. ## Where contracts go wrong The first failure is writing the contract after the code and calling it documentation. That produces an accurate description of whatever the implementation happened to do, including its accidents, and none of the design pressure. The second is over-specifying. A contract that dictates internal structure, database columns or implementation choices is not a contract, it is a design document wearing a costume, and it locks the implementer into the reviewer's first idea. The contract should describe behavior from outside and nothing else. The third is letting it rot. A contract that is not enforced by tests drifts within a month. If the schema does not validate requests at runtime and the CI does not compare the implementation to the spec, the contract is a wish. Write the interface first. Argue about it for an hour. Get a consumer to read it. Generate the tests from it. Then write the code, which will be smaller and clearer than it would have been, because the hard decisions are already made. --- ## A security review you can do in an afternoon 2026-01-17 · audit, security Most teams do not skip security reviews because they do not care. They skip them because they imagine a week of specialists and a report in a language nobody speaks. That version exists and has its place. But the majority of the breaches I have seen in audits would have been caught by a focused afternoon with the codebase, a browser and a database client. Here is the afternoon. ## The checklist - Authentication on every route: list every endpoint, mark which ones require a logged-in user, and try the unmarked ones without a session - Authorization on every object: log in as one user, take an ID that belongs to another, and request it, in every endpoint that takes an ID - Secrets: search the repository history, environment files, logs and error responses for keys, tokens and passwords - Input at the boundary: find every place user input reaches a query, a shell, a file path or an HTML template, and check what validates it - Dependencies: run the audit tool for your package manager and read the critical findings, not the count - Headers and cookies: check that session cookies are HttpOnly and Secure, and that the basic security headers are set - Rate limits: find the login, password reset and any endpoint that sends email or money, and check whether anything limits how often they can be called ## Where the afternoon actually goes Two of those seven checks will consume most of the time and produce most of the findings: authorization on every object, and input at the boundary. The first is tedious and mechanical, which is exactly why nobody does it. Take a real ID from user A, log in as user B, and hit every endpoint that accepts an ID: view, edit, delete, export, download. In the majority of systems I audit, at least one endpoint checks that you are logged in and forgets to check that the thing you asked for is yours. That single class of bug, an insecure direct object reference, has exposed more customer data than every clever exploit combined. The second check is reading. Follow each user-controlled value from where it enters to where it is used. If it reaches a SQL string by concatenation, a shell command, a file path, or a template without escaping, you have found something. Parameterized queries and a templating engine that escapes by default eliminate most of this, and it takes twenty minutes to confirm whether the codebase uses them consistently or 'mostly'. ## What you will find in the logs The secrets check surprises teams most. Search the logs for 'password', 'token', 'authorization' and 'card'. Then search the git history, not just the current tree, for the same. Then trigger an error on purpose and read the response the client receives. I have found full database connection strings in stack traces returned to the browser, and API keys committed in the first week of a project and 'removed' in the second, still sitting in history for anyone with clone access. ## What this is not This is not a penetration test and it will not find a subtle cryptographic flaw or a race condition in the session handling. It is a floor, not a ceiling. But it is a floor that most systems have never been checked against, and the difference between a system that has been through this afternoon and one that has not is the difference between an attacker needing skill and an attacker needing a browser. ## Make it recurring The afternoon is worth more if it happens every quarter, or on every major feature. Write the seven checks into a document, assign one person, put it in the calendar. A team that does this will still have security problems. They will not have the embarrassing ones, and the embarrassing ones are the ones that end up in the news. --- ## WebSockets, SSE or polling: a decision table 2026-01-15 · fullstack, system-design Every product eventually needs something on the screen to update without a refresh. A notification count, an order status, a chat message, a dashboard number. At that moment someone says WebSockets, because it is the answer they have heard of, and the team inherits a persistent connection layer they did not need. There are three honest options, and the right one falls out of four questions: which direction does data flow, how often does it change, how many clients are connected, and what can your hosting keep alive. ## The three options, plainly Polling is the client asking on a timer. It is stateless, it works through every proxy and cache in existence, and it costs one request per client per interval whether or not anything changed. With a fifteen-second interval and ten thousand clients, that is forty thousand requests a minute for data that changes twice an hour. Server-Sent Events is a single long-lived HTTP response that the server writes to whenever there is something new. It flows one way, server to client. The browser reconnects automatically and sends the id of the last event it saw, so the server can resume. It multiplexes fine over HTTP/2 and it is plain text over plain HTTP, which means load balancers and CDNs understand it. In Next.js a route handler can return a readable stream and it works, provided the deployment does not cut long responses. WebSockets is a full-duplex socket upgraded from HTTP. Both sides can send at any time, with low overhead per message. It is the only option when the client needs to push frequently, such as collaborative editing or a game. It is also a stateful connection that a serverless function cannot hold, so it needs a long-running process or a managed service, and that is the cost people forget. ## The decision table - Data changes rarely and freshness within a minute is fine: poll, with a sensible interval and a cache header so the CDN absorbs the load. - Server pushes updates and the client only reads: SSE. Notifications, progress bars, status feeds, streamed model output. - Both sides send often, latency matters, message rate is high: WebSockets, on infrastructure that can hold connections. - Serverless hosting with no persistent process: SSE if the platform allows long responses, otherwise polling. WebSockets need a separate service. - Tens of thousands of concurrent clients: whichever you pick, the connection count is the capacity limit, and a pub/sub layer behind it is mandatory. ## The failure modes Polling fails by cost and by thundering herd. Every client that loaded the page in the same minute polls in the same second. Add jitter to the interval, and back off when the tab is hidden. SSE fails by proxies that buffer. A reverse proxy configured to buffer responses will hold the stream until it fills a buffer, and the client sees nothing for a minute and then everything at once. Disable buffering for the streaming route, send a heartbeat comment every fifteen to thirty seconds so idle connections are not closed, and keep the connection count per process in mind, because each one is an open file. WebSockets fail by state. A connection lives on one server. When you have two servers, a message published on the first must reach clients attached to the second, which means a broker between them. When a server restarts, every client reconnects at once, and if the reconnection has no backoff, the restart becomes an outage. And the connection outlives the session: the user logs out, the socket is still open and still authorized until you close it on purpose. ## What I usually pick For most products, SSE for server-to-client updates and plain HTTP requests for client-to-server actions covers everything. The client sends a mutation the normal way, through a Server Action or a route, the server writes the result and publishes an event, and the event flows down the stream. Two simple mechanisms, no duplex state, and it runs on infrastructure you already have. WebSockets earn their place in the narrow band where the client is a firehose. Everywhere else, they are a persistent connection you are paying to keep alive so that it can carry a notification once an hour. --- ## The outbox pattern in plain words 2026-01-13 · system-design, data Here is a bug that exists in most systems that publish events, and that most teams discover only after it has corrupted something. ## The dual-write problem An order service saves an order to Postgres, then publishes an OrderCreated event to a queue so that billing, email and analytics can react. Two writes, two systems. Now think about what happens between them. The database commit succeeds, then the process crashes, or the broker is unreachable for a second. The order exists, the event was never sent. Billing never charges. Or flip the order: publish first, then commit. The event goes out, the commit fails on a constraint. Billing charges for an order that does not exist. Wrapping both in a try/catch does not help; there is no transaction that spans a database and a message broker. Retrying the publish helps a little and creates duplicates. Distributed transactions across the two exist in theory and are miserable in practice. This is the dual-write problem, and it cannot be fixed by being careful. It has to be designed away. ## Write the event where you write the data The outbox pattern is the design. Instead of publishing to the broker, you insert the event into a table called outbox, in the same database, inside the same transaction as the order. One commit, two rows: the order and the event. Either both exist or neither does. The dual write is gone, because there is only one write. The outbox row carries an id, the aggregate it belongs to (order 123), the event type, the payload as JSON, a created timestamp and a published flag or timestamp. That is the whole table. ## The relay Something still has to move events from the table to the broker. That is the relay: a small process that polls the outbox for unpublished rows, publishes each one, then marks it as published. If the relay crashes after publishing but before marking, it will publish that event again on restart. That is by design, and it leads to the rule that makes the pattern work. Delivery is at-least-once. Every consumer must be idempotent. Billing must be able to receive OrderCreated for order 123 twice and charge once, which usually means storing the event ids it has processed and skipping repeats, or making the operation naturally safe to repeat. This is not optional; a consumer that is not idempotent will eventually double-charge someone. Ordering: a single relay reading rows in insertion order and publishing to a partition keyed by aggregate id gives you in-order delivery per order. Events for different orders can interleave freely, and that is fine. Do not promise global ordering; you do not need it and it does not scale. ## Cleanup and the CDC alternative The outbox grows forever unless you delete published rows. A job that removes rows older than a few days is enough; keep the window long enough to debug an incident. Index on (published, created) so the relay's poll stays cheap even when the table is large. Polling adds latency, typically a few hundred milliseconds to a couple of seconds. If that matters, the alternative is change data capture: a tool like Debezium reads the database's write-ahead log and publishes outbox inserts as they commit, with no polling and no extra load on the table. Same pattern, same table, different relay. CDC is more infrastructure to run, so start with polling unless you have measured that you need better. ## When to bother If a lost or duplicated event only affects a dashboard, you may not need this. If it affects money, inventory, a notification a person is waiting for, or anything that must reconcile later, you do. The outbox costs one table and one small process, and it converts a class of bugs that are nearly impossible to reproduce into a class that does not exist. --- ## Research is a system with a throughput problem 2026-01-10 · science, system-design When I look at a research group the way I look at a production system, I do not see slow people. I see a pipeline with a queue at every handoff and nobody measuring the queues. The thinking is fast. Everything between one thought and the next is where the months go. ## Draw the pipeline Take any study and write down the stages: question, design, approval, data collection, cleaning, analysis, writing, internal review, submission, external review, revision, publication, and finally someone else reading it and building on it. Now write next to each stage how long the work actually sits there waiting versus how long someone is actively working on it. In every lab I have talked to, the waiting dominates. The analysis takes days. Waiting for access to the data takes weeks. Writing takes weeks. Waiting for review takes months. This is the same picture I draw for a deploy pipeline. The compile takes forty seconds. The pull request waits three days for review. Nobody optimizes the compiler when the review queue is the bottleneck, and yet in science we keep asking researchers to think faster. ## Little's law applies to labs There is a simple relationship from queueing theory: the average number of things in progress equals the arrival rate times the average time each thing spends in the system. Fix the rate at which a lab can actually finish things, and every extra project you start increases the time every project takes. Labs run with enormous work in progress. Five studies at different stages, each waiting on something, each context-switched into and out of. The result is that everything is late and nothing is abandoned. The engineering answer is not motivational. It is to cap work in progress, make waiting visible, and attack the largest queue first. When teams do this in software, delivery times drop without anybody working harder. I have watched it in dozens of projects. There is no reason the same physics stops at the door of a laboratory. ## The handoffs lose information Throughput is not only about time. Every handoff in a research pipeline is also a lossy channel. The raw data gets cleaned in a spreadsheet nobody versioned. The analysis lives in a notebook that ran once on a laptop that has since been replaced. The methods section compresses six months of decisions into three paragraphs. By the time a result is published, the path from observation to claim has been reconstructed from memory at least twice. In a software system I would call this a lack of provenance and treat it as a defect, because it is one. It makes replication expensive, it makes reuse nearly impossible, and it means the next person in line starts from zero. A pipeline that loses its intermediate state has a throughput problem even when it looks busy. ## What the fix looks like I am not proposing to turn science into a factory. The creative part, the actual thinking, should stay slow and human and unmanaged. What should not stay slow is everything that is not thinking: finding the relevant prior work, getting a dataset into a usable shape, rerunning an analysis with one parameter changed, producing a figure that matches the data, tracking which version of what produced which number. Those are engineering problems, and they respond to engineering: stable interfaces, automation, versioning, visibility of where work is sitting. The reason I care about this as a systems architect is that a small improvement in throughput compounds. A lab that finishes studies in half the time does not do twice the science. It asks better questions, because the cost of asking a question fell. The scientists I talk to already know where their queues are. They just never had anyone treat it as a system worth designing. --- ## Blue-green, canary and the honest rollback 2026-01-08 · infra Deployment strategies are usually explained as a menu. Rolling update, blue-green, canary, pick one. I think that framing is backwards. The strategy is not the point. The point is what you can do when the new version is bad, and how fast, and with how much data damage. Choose the rollback first, then the strategy falls out of it. ## What each strategy actually gives you A rolling update replaces instances one at a time. It is cheap and it is the default almost everywhere. Its rollback is another rolling update in the opposite direction, which takes as long as the deploy did. During that window, some users hit the bad version. For a service with a handful of instances and a fast startup, this is fine. For a slow-starting service with thirty instances, a bad deploy means a bad quarter hour. Blue-green runs two complete environments and flips traffic between them at the load balancer or DNS. Rollback is the flip reversed, and it takes seconds. The cost is double the compute during the transition and, more importantly, a shared database that both environments must agree on. Blue-green solves the code rollback beautifully and does nothing at all for the data rollback. Canary sends a small share of traffic, say one to five percent, to the new version, watches error rate and latency for a fixed period, then widens. Rollback during the canary phase affects only the canary slice. This is the strategy that turns a deploy into an experiment with a blast radius you chose. It requires that your metrics can be split by version, which is a tagging problem, not a platform problem. ## The dishonest rollback Here is where most deployment plans fall apart. All three strategies roll back code. None of them roll back state. If the new version ran a migration that dropped a column, the old version will crash on startup. If the new version wrote rows in a new shape, the old version will misread them. If the new version published events with a new schema, the old consumers have already rejected them or, worse, processed them wrongly. The honest rollback is one where you have already asked: what has the new version changed that the old version cannot understand? If the answer is 'nothing', you can roll back with any strategy. If the answer is 'the schema', you cannot roll back at all, no matter how green your blue is. So the discipline is: schema changes are additive and shipped alone; the code that uses them ships in a later release; the cleanup that removes old columns ships in a third. Every release in that chain is individually reversible. It feels slow. It is the only way I know to keep 'we can roll back' true. ## What I recommend for a small team For most products with modest traffic, I recommend a rolling deploy with health checks that gate each instance, plus a manual canary for the risky releases. The manual canary is nothing more than deploying to one instance, or one region, or one tenant, and watching for fifteen minutes before continuing. It does not need a service mesh. It needs a version label on the metrics and a person willing to wait. Blue-green earns its cost when startup is slow, when the release is large, or when the business cannot tolerate even a minute of mixed versions. A full canary with automatic analysis earns its cost when deploys are frequent enough that humans cannot watch every one. ## The checklist before every release Can two versions run at the same time against the same database? Is every migration in this release backward compatible with the previous code? If we revert in ten minutes, what data will have been written in a shape the old code cannot read? Is there a metric, split by version, that will tell us within five minutes that the new version is worse? If any answer is 'no' or 'we are not sure', the strategy does not matter. You are not rolling back. You are rolling forward with a fix, under pressure, and you should plan the release as if that were true, because it is. --- ## Hexagonal architecture without the ceremony 2026-01-06 · architecture Hexagonal architecture has one idea worth keeping: the core of your application should not know what database, framework, queue or HTTP library it is running on. Everything else that usually comes with it, the six-layer folder structure, the interfaces for everything, the mapper classes between identical DTOs, is ceremony that teams adopt because a diagram told them to. I have audited codebases where a single 'create order' operation passed through a controller, a request DTO, a use case interface, a use case implementation, a domain service, a repository interface, a repository implementation, an entity mapper and an ORM model. Nine files. The business rule was four lines. The ceremony had eaten the architecture. ## The one rule that matters Dependencies point inward. The domain logic imports nothing from the outside world. The outside world imports the domain. That is the entire principle. If your pricing rules can be unit tested without a database, a framework or a network, you are doing hexagonal architecture, regardless of what your folders are called. Everything else is a technique for enforcing that rule, and techniques should be chosen by cost. A TypeScript project can enforce it with a lint rule that forbids imports from 'infra' inside 'domain'. That is one line of configuration. It does not require an interface for every repository. ## What I actually do in a Next.js project Three folders per feature, not six. A domain folder with plain functions and types: no classes required, no decorators, no framework imports. This is where the rules live, and it is the only folder that gets thorough unit tests. An application folder with the use cases, which are functions that take the domain plus a small set of injected capabilities: a way to load an order, a way to save it, a way to publish an event. And an infra folder where those capabilities are implemented against Postgres, a queue, an email provider. The injected capabilities are the ports. They are just TypeScript function types, defined next to the use case that needs them. Not an abstract class in a separate package. Not a generic Repository that every entity inherits from. A use case that needs to load and save orders declares exactly those two functions and nothing more. The adapter is a file in infra that exports an object matching that type. Server Actions and route handlers are the driving adapters. They parse the request, call a use case with the real infra wired in, and shape the response. They contain no rules. If a route handler has an if statement about business logic, it is in the wrong place. ## The ceremony I refuse Interfaces with a single implementation, created 'in case we swap the database'. In fifteen years I have seen the database swapped maybe three times, and each time the interface did not survive the swap anyway because the new database had different transaction semantics. Write the interface when the second implementation arrives, which is usually the in-memory fake for tests. That one is worth it. DTO-to-entity-to-model mappers when the three shapes are identical. Mapping earns its keep when shapes actually differ, when the wire format has a field the domain must not see, or when the domain has a computed value the database does not store. Mapping identical objects through three layers is noise that hides the two places where it matters. A separate 'domain' package published to a registry so it can be 'reused'. It never is. It just adds a build step and a version number to every change. Anemic entities with a service layer that does all the work. If the domain is only data and the rules live in services, the 'domain' folder is a folder of types and the architecture is procedural with extra steps. Either put the rules on the domain objects or admit you are writing procedural code, which is fine, and stop pretending. ## How to tell if it is working The domain folder has the highest test coverage and the fastest tests. A new adapter, say switching from one email provider to another, touches one file in infra and zero files elsewhere. A new business rule touches the domain and maybe one use case, and never a route handler. And a new engineer can find where 'refund eligibility' is decided in under a minute by looking at file names. If those four things are true, keep whatever structure got you there. If they are not, no number of additional folders will fix it. Hexagonal architecture is a direction of dependencies. It is not a folder layout. --- ## The generalist advantage 2026-01-05 · career, fullstack I am a full-stack engineer in the widest sense of the term: I design the system, write the frontend, write the backend, set up the infrastructure, and then audit the whole thing when it fails. For most of my career this was treated as a weakness. Specialists were the serious ones. The generalist was the person who knew a little about everything and not enough about anything. I no longer think that is true, and I think the reason is structural, not personal. ## Failures live at the boundaries When I audit a system, the failure is almost never inside a well-defined component. The frontend team's code works. The backend team's code works. The database is fine. The failure is in the space between them: the assumption the frontend made about a field the backend later changed, the retry the client does that the server was not designed to absorb, the cache that was correct until someone added a second writer. A specialist sees their component. A generalist sees the seam. Since the seam is where the failures live, the generalist has an advantage in exactly the work that matters most when things go wrong. This is a large part of why I find failures other people miss. Not because I know any single layer better than its specialist, but because I can follow a request across all of them without handing it off. ## Systems thinking requires seeing the whole system System Design as a discipline is about trade-offs across layers. Do we validate here or there. Do we push this to the client or keep it on the server. Is this a database constraint or an application rule. Each of these questions is unanswerable from inside one layer, because the cost of the decision lands in a different layer from the one where it is made. You cannot do System Design well as a pure specialist. You can contribute to it. But the person holding the whole picture must have worked in every part of it, at least enough to know what each part costs. ## What the generalist must do to stay credible The criticism of generalists is not entirely wrong. A generalist who is shallow everywhere is not useful. So the discipline is this: pick one or two areas to be genuinely deep in, and be competent in the rest. Deep means you could teach it. Competent means you could ship it, review it, and know when to call the specialist. For me the deep areas are System Design and finding failures. The competent areas are everything else, with a preference for Next.js and TypeScript because they let me keep a single mental model from the browser to the database. The stack choice is in service of the whole picture, not the other way round. ## The career shape A specialist career is a ladder inside one column. A generalist career is a path across columns, and it is harder to describe on a résumé. Nobody has a job title for "person who can see the whole thing". What I have found is that the generalist path gets more valuable with age, not less. Early on, the specialist is easier to hire because the role is clear. Later, when the company has five specialists who each say the problem is in someone else's layer, the person who can walk the whole request through and point at the seam becomes the most valuable engineer in the building. That is the position I have built my studio around, and it is the position I would tell any engineer with broad curiosity to aim for. ## The practical version If you feel the pull toward breadth and have been told to narrow down, here is a different instruction. Go deep in one thing until you could teach it. Then, instead of going deep in a second thing, go across. Build one project end to end by yourself, including the infrastructure, including the part you are bad at. Do it again. After the third time you will see seams that specialists cannot see, and that is not a consolation prize. It is the advantage. --- ## Embeddings are not magic. They are a compression of your data. 2026-01-03 · ai, data Every RAG proposal I review has the same slide: documents go in, embeddings come out, similarity search finds the answer. The arrow between chunk and vector is drawn as if it were free and exact. It is neither. An embedding is a lossy compression of a piece of text into a few hundred or a few thousand numbers, and the thing deciding what gets kept and what gets thrown away is a model trained on someone else's data for someone else's purpose. Once you see it that way, most retrieval failures stop being mysterious. ## What gets thrown away A fixed-size vector cannot hold everything in a paragraph. The embedding model keeps what was useful for its training objective, which is usually some form of "do these two texts talk about the same thing". Topic survives. Tone mostly survives. Exact identifiers, numbers, negations and dates survive badly. Two chunks that say "the patient was given the drug" and "the patient was not given the drug" land close together in most embedding spaces, because they are about the same thing. In a system where that distinction matters, and in mine it always does, cosine similarity is the wrong tool for the last mile. I have audited pipelines that returned the right document with the wrong polarity and nobody noticed because the retrieval metrics looked great. ## Your data is not the training data Embedding models are trained on general text. Your corpus is contracts, or tickets, or lab notes, or product SKUs. When the vocabulary of your domain is rare in the training set, the model compresses it poorly: internal acronyms, product codes and part numbers all collapse toward a generic "technical noise" region. The symptom is retrieval that works beautifully on the demo queries written in plain English and fails on the queries real users type, which are full of the identifiers the model never learned. The fix is rarely a bigger embedding model. It is hybrid retrieval: keyword or BM25 for the exact tokens, vectors for the meaning, and a reranker to combine them. ## Chunking is a compression decision The chunk is the unit of compression, and nobody talks about it. A chunk of two thousand tokens embedded into one vector is an average of everything inside it, so a specific fact in the middle is diluted by the boilerplate around it. A chunk of fifty tokens preserves the fact and loses the context that tells you what it refers to. There is no universal right size. There is a right size for your documents and your queries, and you find it by measuring, not by copying the default from a tutorial. I keep a small eval set of real queries with known correct passages and I rerun it every time the chunker, the model or the corpus changes. That set has caught more regressions than any dashboard. ## Treat it like a format, not an oracle - Version the embedding model like a schema, because changing it invalidates every vector you stored. - Store the source text next to the vector, always, so you can rerank and display what was actually retrieved. - Measure recall on your own queries, not on a public benchmark. - Combine vectors with exact matching for anything that has an identifier. When a team internalizes that an embedding is a compression format with a particular loss profile, the architecture conversation changes. They stop asking "which vector database" and start asking "what does my retrieval need to preserve, and what am I allowed to lose". That is the right question, and it has the same shape as every other data engineering question I have ever answered. --- ## Server Components change where the boundary is 2025-12-30 · fullstack, nextjs, architecture For fifteen years the architecture of a web application had one obvious boundary: the API. The browser was on one side, the server on the other, and JSON crossed between them. Every team drew that line the same way because the technology forced it. React Server Components erase that line and draw a new one. Most teams I audit have adopted the new tools without noticing that the boundary moved, and that is where their problems come from. ## The boundary is now a directive In the App Router, every component is a Server Component unless a file starts with `'use client'`. Server Components run only on the server. They can read the database, read the file system, use secrets, and none of their code ships to the browser. Client Components run in both places, first on the server for the initial render and then in the browser. So the architectural boundary is no longer a URL. It is the point where a Server Component renders a Client Component and passes it props. Everything crossing that point must be serializable. Functions cannot cross, except Server Actions, which are references to server code. Class instances cannot cross. Dates become strings unless you handle them. That is a real boundary with real rules, and it deserves the same care you used to give an API contract. The difference is that nobody writes a spec for it, because it looks like a prop. ## What moves to the server The most useful consequence is that data fetching moves up. A page component awaits its queries directly. Child Server Components can fetch their own data too, and request-level memoization means the same fetch in three components is one network call. No client state library, no loading spinner choreography, no waterfall of useEffect calls. It also moves dependency weight to the server. A markdown renderer, a date formatting library, a syntax highlighter: if only a Server Component imports it, the browser never downloads it. I have cut client bundles by half in audits just by moving imports across the boundary. And it moves authorization to the server by default. If a component that renders admin data is a Server Component, the check happens where it cannot be bypassed by editing JavaScript in the browser. That is not a feature. It is the removal of a whole class of vulnerability. ## The failures I find The first failure is the accidental client tree. Someone adds `'use client'` to a layout for a menu toggle, and now every component under that layout is a Client Component, including the ones that were supposed to fetch data on the server. The fix is composition: keep the interactive part small and pass Server Components into it as children. The second is leaking. A Server Component fetches a full user record and passes the whole object to a Client Component that only needs the name. The password hash is now in the HTML payload. I check for this in every audit, and I find it more often than I should. Pass the fields you need, nothing else, and use the `server-only` import in any module that touches secrets. The third is treating a Server Action as a private function. It is a POST endpoint that anyone can call with any payload. Validate the input, check the session inside the action, and never trust an id that came from the client to be one the user is allowed to touch. ## Drawing the line on purpose My rule: a component becomes a Client Component when it needs an event handler, browser state or a browser API. Nothing else earns the directive. The line goes as low in the tree as possible, and the code that crosses it is the thinnest data I can get away with. The boundary did not disappear. It moved to a place that is easier to cross and easier to cross badly. Treat it like the API it replaced. --- ## The staging environment lie 2025-12-27 · infra, engineering Every team has a staging environment, and every team believes it more than it deserves. 'It passed staging' is offered as evidence that a release is safe. What it actually proves is that the code starts, the happy path works, and nothing obvious is broken when one person clicks through it. Those are useful facts. They are not the facts that prevent outages. The lie is not that staging is useless. The lie is that staging resembles production. In the systems I audit, it almost never does, in the ways that matter. ## The ways staging differs Data is the first and largest gap. Staging has a thousand users and production has a million. The query that runs in ten milliseconds on staging takes eleven seconds in production because the planner chose a different strategy at scale. The migration that took two seconds on staging locks the table for three minutes in production. No amount of clicking through staging finds either problem. Traffic is the second. Staging serves one tester at a time. Production serves hundreds concurrently, and concurrency is where the race conditions, the lock contention, the connection pool exhaustion and the cache stampedes live. Code that has never run under concurrent load has never been tested for its most common production failure mode. Configuration is the third. Staging shares a database instance to save money, uses a smaller cache, points at sandbox versions of third-party APIs that behave differently from the real ones, has different timeouts, different feature flags, a different CDN policy or none at all. Every one of those differences is a bug that staging cannot catch by construction. And there is drift. Staging was identical to production on the day it was created. Since then, someone tested a config change on staging and never reverted it. Someone else changed production by hand during an incident. The two environments have been diverging silently for months, and nobody has a diff. ## What staging is actually for Staging is good at three things. Catching integration problems between services that were developed separately. Letting a non-engineer see a feature before it ships. Running the automated end-to-end tests against a deployed system rather than a mocked one. Those are real benefits, and a team should keep staging for them. What staging is not for is answering 'will this survive production'. That question has other answers, and they are cheaper than a second production. ## The alternatives that work - Test migrations against a recent copy of the production database, restored to a throwaway instance, and time them. This finds the lock problems that staging's small tables hide. - Load test the specific endpoint that changed, at production concurrency, before merging. Not the whole system, just the path that is new. - Deploy to production behind a feature flag and turn it on for internal users first. This is staging with real data, real traffic and real configuration. - Canary the release to a small fraction of real traffic and watch error rate and latency by version for fifteen minutes. - Keep infrastructure as code for both environments from the same source, and run a diff regularly to catch drift. Every one of those tests the code against the thing that will actually run it. Staging tests the code against a smaller, older, quieter imitation. ## Where to spend the money Teams often spend more on making staging bigger than they would spend on the alternatives above. A staging environment that is a faithful copy of production, with production-sized data, production traffic replay and identical configuration, is a second production, and it costs like one. Very few teams need that. What they need is the discipline to test data changes against real data sizes, code changes under real concurrency, and releases against real traffic in small doses. Keep staging. Just stop asking it questions it cannot answer, and stop treating 'it passed staging' as a reason to skip the questions that matter. --- ## What a new service really costs 2025-12-23 · architecture, system-design A new service on the architecture diagram costs about a day to draw and about a week to scaffold. That is the part everybody estimates. The part nobody estimates is what it costs for the next five years, and that number is what I want to make visible, because it changes decisions. I am not against services. I am against services that were created because a box looked cleaner than a module, by people who never saw the bill. ## The fixed costs Every service, regardless of size, carries a floor of cost that a module inside an existing deploy unit does not. It needs a repository or a workspace entry, a CI pipeline, a container image, a deploy pipeline with at least two environments, a health check, a way to get secrets, a log shipper, a metrics exporter, trace context propagation, alerts, a dashboard and a runbook. Each of those is a few hours the first time and then permanent maintenance: the base image goes out of date, the CI runner changes, the logging library has a vulnerability. It needs an identity: a service account, network rules, permissions to the things it talks to. Someone must decide those permissions and someone must review them when they change. And it needs ownership. A named team, an on-call rotation that knows it exists, a place in the incident process. Services without a clear owner are the ones that stay on an unpatched runtime for three years. My rough number: a service costs a minimum of two to four engineer-days per quarter just to exist, doing nothing, before any feature work. Ten services is a month per quarter of pure upkeep. Fifty is a full-time engineer whose job is keeping boxes alive. ## The costs that scale with connections The fixed cost is the small one. The bigger one grows with the number of things the service talks to. Every call between services is a network call that can fail, time out, or return a version you did not expect. Each of those needs a timeout, a retry policy, a circuit breaker decision, an idempotency story and a contract test. A function call inside a monolith needs none of that. Every piece of data that two services both need is now either duplicated, with a sync problem, or fetched at runtime, with a latency and availability problem. A join that was one SQL line is now two calls and a merge in memory, or an event pipeline that keeps a copy fresh. Every transaction that crosses a boundary is now a saga, or it is quietly not a transaction anymore. I have audited many systems where an operation that should have been atomic was split across three services, and the failure mode was money in an inconsistent state that a nightly job 'fixed'. Debugging follows the same curve. A stack trace becomes a trace across services, which is fine if tracing is set up everywhere and terrible if one hop drops the context. The first hour of every incident goes to finding which box is actually failing. ## The costs on people There is a quieter cost. Every service is a place a new engineer must learn to find. It is a repository they must clone, a local setup they must get working, an owner they must ask. Onboarding time scales with service count more than with lines of code. And every boundary is a place where 'not my service' can be said. Architectures with many small services tend toward many small blind spots, where a bug lives on the seam and neither side owns the seam. ## When the cost is worth it The bill is worth paying when the service buys something a module cannot. Independent scaling, when one part of the system genuinely has a different load profile and you have measured it. Independent deploys, when two teams are blocked on each other's releases and you have counted the blocks. Isolation, when a component is risky, in a different language, or must be sandboxed. A real ownership boundary, when a separate team will own it for years. If none of those apply, the honest choice is a module with a clear interface inside the existing deploy unit. It gets you most of the design benefits, the separation, the named boundary, the contract, at almost none of the operational cost. And when one of the reasons above becomes true, the module is already shaped to be extracted. So my rule before adding a box: write the invoice. Fixed costs, connection costs, people costs, in a paragraph. Then write what the box buys. If the second paragraph is shorter than the first, draw a module instead. --- ## The model will change. The interface should not. 2025-12-20 · ai, architecture In the time I have been shipping features on language models, every model I started with has been replaced. Deprecated, superseded, outpriced or simply outperformed. This is not going to stop. If your product's behavior is wired directly to one provider's SDK, one prompt format and one set of quirks, every model change is a migration project. The way out is the same one systems engineers have used for databases and payment providers for decades: put an interface in front of the thing that changes, and make the rest of the system depend on the interface. ## What the interface owns The interface I put around a model is narrower than most SDKs. It takes a typed request: a task name, structured inputs, an output schema, and constraints like max latency and max cost. It returns a typed result, validated against the schema, with metadata: which model answered, tokens used, latency, and whether a fallback was used. Prompt templates live behind the interface, keyed by task, versioned. Provider clients live behind it. Retries, timeouts and rate limits live behind it. The application never sees a raw completion string and never builds a prompt. It asks for a task to be done and gets a typed answer or a typed error. ## Why tasks, not prompts Naming the unit "task" instead of "prompt" is the most useful decision in this design. A task like "classify support ticket" has a stable contract: these inputs, this output shape, this quality bar. How it is fulfilled is an implementation detail. Today it may be a frontier model with a long prompt. Next quarter it may be a small model fine-tuned on your own data, or a cached lookup for the ninety percent of tickets that repeat. The application code does not change. The eval set for that task does not change either, which means you can measure whether the new implementation is as good as the old one before you flip the switch. ## The eval set is the contract test An interface without a test is a hope. For every task behind the boundary I keep an eval set: real inputs, expected outputs or a scoring rule, and a threshold. Swapping the model means running the set against the new implementation and comparing. This is exactly what a contract test does for an HTTP API, and it is the only thing that makes model changes boring. Without it, every upgrade is a week of manual spot checks and a silent regression somewhere nobody looked. ## Things that leak - Provider-specific features like tool-calling formats or special tokens, used directly in application code. - Output parsing that assumes a particular model's formatting habits. - Prompts that reference the model by name or rely on one model's tolerance for ambiguous instructions. - Cost assumptions baked into product decisions, like "we can afford to call it on every keystroke". Each of these is a place where the model has escaped the boundary and become part of the product. I look for them in audits, and I find them in almost every codebase that grew a model feature quickly. ## Routing is the payoff Once the interface exists, routing becomes cheap. Send the easy inputs to a small model and the hard ones to a frontier model. Send a percentage of traffic to a candidate and compare. Fall back to a cached answer when the provider is down. None of this is possible when the application talks to a provider directly, and all of it is a few lines when it talks to an interface. The model is the fastest-changing dependency you have. Treat it like one. --- ## Why I still write SQL by hand 2025-12-18 · fullstack, data I use an ORM in every product I build. I also write SQL by hand in every product I build, and the second habit surprises people more than it should. The two are not in conflict. They cover different parts of the same job, and the part I write by hand is the part where the product's performance and correctness actually live. The database is the most capable engine in the stack. It can sort, aggregate, window, lock, deduplicate and paginate millions of rows in milliseconds, with decades of engineering behind each operation. A query builder exposes a fraction of that. For the queries that matter, I want all of it. ## The queries I always write myself Reporting and aggregation. Anything with a group by, a window function, a common table expression or a lateral join. Builders can express some of these, but the result reads like a puzzle and the generated SQL is often worse than what the planner would like. I write the query, read the plan, adjust, and keep it as a named function with a typed result. Keyset pagination. Offset pagination is what every builder gives you and it degrades linearly with page number, because the database still walks every skipped row. Keyset pagination, where the query says 'rows after this id and timestamp', is constant time at any depth and needs a hand-written where clause with a composite comparison. Any list a user can scroll deep into gets this. Upserts and conditional writes. Insert on conflict update, with a where clause that only applies the update if the incoming version is newer. That is idempotency and concurrency control in one statement, atomic, and most builders either cannot express the condition or express it badly. Queue operations. Select for update skip locked is how you build a reliable job queue on a relational database without a broker, and I have replaced several message brokers with that one line. No builder I have used generates it cleanly. Bulk operations. Updating a hundred thousand rows through an ORM is a hundred thousand objects and a hundred thousand statements. A single update with a join from a values list is one statement and a fraction of the time. ## Keeping it safe and typed Hand-written SQL has a reputation for injection, and that reputation belongs to string concatenation, not to SQL. Every query I write uses parameters. The tagged template style, where the values are interpolated as placeholders by the library rather than as text, makes the safe way the easy way and the unsafe way visibly different. The result is typed. Either the tool generates a TypeScript type from the query by asking the database what columns it returns, or I declare the row type and validate the first row in tests against a real database. Both are fine. What is not fine is a query returning any that gets asserted into a shape at the call site. Queries live in modules named for the aggregate they serve, next to the ORM code for that aggregate, and nothing else in the codebase touches a table directly. The domain layer calls functions. It does not know which ones are SQL and which are the ORM, and that is the point. ## What writing SQL teaches you The real reason I keep the habit is that it keeps me honest about what the database is doing. When you write the join yourself, you think about the index. When you write the aggregation, you think about the row count. When you write the lock, you think about the transaction boundary. The ORM lets you not think about those things, which is a feature right up until it is the incident. Every audit where I found a slow database, the fix started with someone reading the SQL for the first time. I would rather that person be the one who wrote it, on the day they wrote it, with the plan open in the next tab. Use the ORM for the ninety percent. Write the ten percent by hand, keep it typed and parameterized, and read the plan. That ten percent is where the product's speed comes from, and it deserves a human author. --- ## A graceful degradation checklist for product teams 2025-12-16 · system-design, reliability Most outages are not total. The database is slow, not down. The recommendations service is timing out. The payment provider is answering in eight seconds instead of one. What the user sees in those moments is a product decision, and it's usually made by accident, by whichever error path a developer wrote first. Graceful degradation is choosing on purpose. This is the checklist I walk through with product teams, ideally before the first incident, and realistically after it. ## Decide what's core Start by sorting features into core and optional. Core is what the user came for: viewing the catalogue, checking out, reading their messages. Optional is everything that improves it: recommendations, personalised ordering, live stock counts, reviews, the chat widget. The rule is that an optional feature failing must never take down a core one. That sounds obvious and is violated constantly: a recommendations call on the product page with no timeout, blocking the render of the product itself. Once you've sorted, each optional feature needs a fallback that a designer has seen. An empty space, a generic list, a 'popular right now' section served from a static cache. Whatever it is, it's not a spinner and it's not an error. ## The checklist - Recommendations down: show a curated or cached list, or hide the block entirely. Never block the page. - Search down: fall back to category browsing and a simpler database query. Say 'search is limited right now' rather than showing zero results. - Payments slow: show progress honestly, keep the button disabled, and never let a double-tap create two charges. Have a threshold after which you tell the user to wait for an email confirmation. - Payments down: accept the order into a pending state if the business allows it, and confirm later. If not, say so clearly and preserve the cart. - Read replica lagging: label the data with 'as of' and a timestamp instead of pretending it's live. - Primary database down or in failover: switch to read-only mode. The user can browse, cannot buy, and is told why. - Any write path failing: queue the write, confirm to the user that it was received, and process when the dependency returns. Only acceptable if 'received' and 'done' are different states in the UI. - Third-party analytics or tracking down: the product must not notice. Fire and forget, with a short timeout. - Feature flags: every optional feature has a kill switch that flips without a deploy, and the person on call knows where it is. ## Stale beats missing, if you label it There's a strong temptation to hide anything that isn't fresh. Resist it. Yesterday's stock level with a label is more useful than an error, as long as the user can tell the difference. The failure mode to avoid is stale data presented as live: a sold-out item shown as available, a paid invoice shown as due. The label is what makes staleness honest. 'Updated 14 minutes ago' costs nothing and prevents a support ticket. The same applies to queued writes. 'We received your order and will confirm by email' is fine. 'Your order is confirmed' when it's actually sitting in a retry queue is not. Match the words to the state. ## Read-only mode is a feature, build it Read-only mode is the single most valuable degradation mode and the least often built. It needs a global flag, a middleware that rejects writes with a clear message, and a UI that hides or disables write actions instead of letting the user fill a form and fail at the end. Build it on a quiet week. During a database failover it turns a full outage into a mild inconvenience. ## Practise it None of this works if the first time you flip the switch is during the incident. Turn recommendations off on a Tuesday afternoon and look at the page. Put the system in read-only for ten minutes in staging with the real UI. Kill the payment provider in a sandbox and watch what a double-click does. Product teams that have seen their degraded states are calm in incidents. The ones that haven't are learning what the product looks like at the same time as their users. --- ## What clinicians taught me about requirements 2025-12-14 · science, career I have gathered requirements for marketplaces, logistics platforms, financial products and factories. I thought I was good at it. Then I started spending time with clinicians, first around a clinic scheduling app years ago, later around research tools, and I learned that everything I knew about requirements had been calibrated to people whose worst case is a refund. ## They describe the exception first Ask a product manager how a feature works and they describe the happy path. Ask a clinician how a process works and they describe the case where it went wrong. The patient who had two records. The result that came back after the decision was already made. The allergy that was in the note but not in the field. It took me a while to understand this is not pessimism. It is training. Medicine teaches people that the normal case takes care of itself and the exception is where someone gets hurt. That is precisely how I audit systems, and I had never seen it as a requirements technique before. Now I open every discovery session with the same question: tell me about the last time this went wrong. ## Interruption is the default state Every workflow I had ever modeled assumed the user starts a task and finishes it. A clinician starts a task, gets called away, comes back forty minutes later, and has to know instantly where they were and whether anything changed while they were gone. A form that loses state on refresh is an annoyance in an e-commerce checkout. In a ward it is a safety issue. Requirements for clinical software have to specify what happens when the task is abandoned halfway, not as an edge case but as the main case. I now ask about interruption in every project, medical or not, because it turns out that warehouse operators and support agents live the same way. Clinicians were just the first to say it out loud. ## They do not trust what they cannot see the source of Show a clinician a number and the first question is where it came from. Not because they distrust the software, but because they have been burned by a value that was copied from an old record, or unit-converted twice, or entered for the wrong patient. A requirement I never used to write, and now write on every project, is that every displayed value must be able to show its origin on demand. The source, the time, the person or process. This is provenance at the interface, and it costs almost nothing if you design for it from the start and a rewrite if you do not. ## Vague requirements are a liability, literally In most software, an ambiguous requirement becomes a bug, then a ticket, then a fix. In medicine, an ambiguous requirement becomes a decision someone made with the wrong information, and there is a name attached to that decision. Clinicians taught me to push back harder on vagueness than I ever did with business clients. 'Show the relevant results' is not a requirement. Relevant to whom, chosen how, and what happens to the ones that were filtered out? ## What I took back to every project The clinical mindset made me a better engineer for non-clinical work. Start with the failure, not the feature. Assume the user will be interrupted. Make every number explain itself. Refuse ambiguity even when it is socially awkward. My father spent forty years in consulting rooms where those habits were the difference between a good day and a tragedy. I never got to learn them from him. I learned them from people who do the same work, and I try to build tools that deserve their trust. --- ## Secrets management for small teams that will grow 2025-12-11 · infra, security Every project starts with a .env file. That is not a mistake. A .env file at the beginning is the correct amount of ceremony for a team of two. The mistake is the day nobody notices that the team is now eight people, the .env file has been pasted into four chat threads, and the production database password is the same one from the first week. I audit systems for a living, and secrets are where I find the failures that nobody else finds. Not because teams are careless. Because the migration from 'small' to 'not small' happened without anyone deciding it had. ## The three stages Stage one is the .env file, gitignored, with a committed .env.example listing every key with an empty value. The example file is the important part. It is the contract that says which secrets exist. If a new secret is added without touching the example, the next developer's setup breaks in a confusing way instead of a clear one. Stage two is the platform's secret store. Every serious hosting platform has one. Secrets are set through a dashboard or a CLI, injected as environment variables at runtime, and never live in a file on anyone's laptop. The transition here is cheap: the code does not change at all, only where process.env gets its values from. Most teams should be here by the time they have a production customer. Stage three is a dedicated secrets manager with versioning, audit logs, automatic rotation and per-service identities. This is where you go when you have multiple services, multiple environments and compliance questions. The code changes a little: secrets are fetched at startup or on demand instead of read from the environment. The mental model does not change if you did stage one and two properly. ## The rules that survive every stage The reason those three stages do not require rewrites is a handful of rules that apply from the first day. Secrets are read in exactly one module, validated at startup, and exposed to the rest of the code as a typed configuration object. Nothing else touches process.env. In a TypeScript project this is a single file with a schema, and the application refuses to boot if a required secret is missing or malformed. That one habit turns a silent 3 a.m. failure into a loud deploy failure. The second rule is that secrets are per environment and per purpose. Development, staging and production never share a credential. The API key for sending email is not the API key for reading analytics, even if the provider would let you use one. When one leaks, and one will, the blast radius is the size of that one purpose. The third rule is that every secret has a known rotation path before it is created. Not a schedule, a path. Which service uses it, where it is set, what breaks during the rotation and in what order to update things. Write this in a comment next to the key in the example file. When the leak happens, you will not be thinking clearly, and the comment will do the thinking for you. ## The things I always find - Secrets committed to the repository at some point in history, still valid, never rotated after being removed from the current tree. - A single shared database user with full privileges used by the application, the migrations, the analytics job and a developer's laptop. - Long-lived tokens for third-party APIs with broad scopes, created during a prototype and still in use. - Secrets printed to logs by a debug statement that serializes the whole config object. - CI pipelines that echo environment variables, with the output retained for months. None of these require a sophisticated attacker. They require one leaked log, one former contractor, one public repository that was private last year. ## What to do this week Scan the git history of every repository for secrets. Rotate everything it finds, regardless of whether you believe it leaked. Move production secrets into the platform store and delete them from every laptop. Split the database user into at least an application role and a migration role. Add the startup validation. That is a few days of work, and it takes you from stage one to a stable stage two without a rewrite. Stage three can wait until the team needs it, and when it does, the shape of the code will already be right. --- ## The database is not your integration layer 2025-12-09 · architecture, data The most expensive integration pattern I find in audits is also the most invisible one. Two applications, sometimes five, all connected to the same database. No API between them. The tables are the contract, and nobody wrote the contract down. It always starts innocently. The reporting tool needs order data, so it gets a read-only user. The new mobile backend needs to update customer profiles, so it gets write access to that one table. Eighteen months later there are seven consumers, three of them writing, and a column rename requires a meeting with four teams and a maintenance window. ## Why it feels free and is not Sharing a database looks free because there is no code to write. The data is already there. Just connect. The cost shows up later, in three specific places. Schema changes become negotiations. The team that owns the orders table cannot add a NOT NULL column, split a table or change an enum without knowing every reader and writer, and they never know all of them. So the schema freezes. The freeze is the real cost: the module that should evolve fastest becomes the one that cannot move. Invariants become unenforceable. The orders application knows that an order in status 'paid' must have a payment row. The reporting script that fixes 'bad data' on Friday afternoons does not. Now the application's own assumptions are violated by writes it never sees, and the bugs appear in the application, where nobody is looking for a cause outside it. Performance becomes shared fate. An analytics query that scans forty million rows locks nothing, in theory. In practice it evicts the working set from cache, saturates I/O and makes checkout slow at the exact moment the finance team runs month-end reports. There is no rate limit, no timeout, no circuit breaker between a SQL client and a table. ## The data owner rule Every table has exactly one application that writes to it. That application is the owner. Everyone else gets the data through something the owner controls: an API, a published event stream, a replicated read model that the owner builds and versions on purpose. This is the whole rule, and it is enforceable with database permissions, which is the part most teams skip. The owner can change its tables freely, because the only consumer of the raw schema is itself. The published shape, the API response or event payload, is versioned and stable. That is the contract. It can be written down, tested and evolved with a deprecation window. ## Getting out without a rewrite You do not need to build a service mesh. The migration I usually recommend has four stages, and each one is independently valuable. - Inventory: turn on query logging or use the database's own connection stats to list every application and every user touching each table. Most teams are surprised by the list. - Freeze writes: revoke write permissions from every non-owner. Give them an endpoint or a queue instead, one table at a time, starting with the table that changes most often. - Replace reads with views or a read model: for consumers that only need data, expose a versioned view or a replicated table that the owner populates. The view is the contract; the underlying tables become private. - Revoke direct access entirely and delete the extra database users. If a consumer still needs raw SQL, that is a signal it needs a proper data product, not a login. The freeze-writes stage delivers most of the safety. Once only one application writes, the invariants are enforceable again and the schema can start moving. Reads can stay on views for a long time without much harm. ## The exception people bring up 'But it is the same team.' Fine. If one team owns both applications and both schemas, the risk is lower. It is not zero. The two applications will still deploy separately, which means there will be a window where one has the new schema expectations and the other does not. That window is where the outage lives. If you truly cannot separate them, at least make every schema change backward compatible and deploy the writer first. But be honest that what you have is one system, and stop calling it two. The database is a fantastic place to store data. It is a terrible place to negotiate between teams. Put the negotiation somewhere you can version. --- ## Read the error handling first 2025-12-08 · audit, reliability If you gave me ten minutes with an unfamiliar codebase and told me to predict its next outage, I would not read the architecture or the main business logic. I would search for every catch, every rescue, every except, every .catch(), and read them in order. Error handling is where a team wrote down, without meaning to, what it feared, what it did not understand, and what it decided to ignore. It is the most honest documentation in the repository. ## The four kinds of catch block The empty catch: an error happened and the system pretends it did not. The log-and-continue: the error is recorded and the function returns as if it succeeded, so the caller proceeds with a lie. The catch-all: anything, including bugs in the handler itself, is caught by the same block and treated the same way. And the rethrow that loses context: the original error is wrapped in a generic one, and by the time it reaches a log, nobody can tell what failed. I count each kind. A codebase with many of the first three is a codebase that has no idea what state it is in after anything goes wrong. ## What a swallowed error costs The damage of an ignored error is rarely where the catch is. It is downstream. A payment provider times out, the catch logs a warning and returns null, the caller treats null as 'not paid', the order is cancelled, and the customer was actually charged. Nothing crashed. Every function did what it was written to do. The catch block made a decision, on behalf of the business, that a timeout means a failed payment, and nobody ever reviewed that decision because it was three lines inside a try. When I read a catch, the question is always: what does the caller believe is true after this line, and is it? If the answer is 'the caller thinks it succeeded', that catch is a bug with a delay on it. ## Retries hide in the same place Right next to the catch is usually the retry, and the retry is usually wrong. No maximum attempts. No backoff, so the failing dependency gets hammered at exactly the moment it is weakest. No idempotency check, so the retry of a write that actually succeeded creates a duplicate. Retrying errors that will never succeed, like a validation failure, alongside ones that might, like a network blip. A retry is a decision to make the same request again, and it deserves the same scrutiny as the original request. Most of the time it gets none. ## The errors that are never caught The mirror image matters just as much. I look for every external call, every parse, every database write, and check whether anything catches its failure. A JSON parse on a webhook body with no try around it. A database call in a background job with no handler, so the job dies silently and the queue moves on. A promise without a catch, which in some runtimes means the process exits and in others means the error vanishes. Unhandled paths are catch blocks that nobody wrote, and the behavior is whatever the runtime decides. ## What good looks like Good error handling is boring and specific. Catch the errors you can actually recover from, by type, and let everything else propagate to a boundary that knows how to fail loudly. At that boundary, return a response that is honest about what happened, record the original error with its context, and leave the system in a state the next request can trust. Every retry has a limit, a backoff and an idempotency key. Every background job reports its own death. None of this is hard. It just requires someone to decide that the failure path is part of the product, and to read it as carefully as the success path. --- ## The cost of a token in production 2025-12-06 · ai, business Every team that ships an LLM feature reads the provider's price list once, multiplies by an imaginary number of requests, and decides it is cheap. Then the first monthly invoice arrives and it is five times the estimate. Nothing was wrong with the price list. What was wrong was the model of how tokens get consumed. A token in production is not a unit price. It is a unit price multiplied by a shape, and the shape is what you actually control. ## Tokens are a unit cost with a shape The bill for a request is input tokens plus output tokens, and the input side is almost always where the surprise lives. A system prompt of two thousand tokens is sent on every call. A conversation history grows with every turn. A RAG context stuffed with ten chunks because ten felt safer than four. A tool-calling loop that sends the whole transcript back on every step. Multiply any of those by real traffic and the per-token price stops mattering. The shape matters: how much you send, how often, and how many times per user action. ## Where the money actually goes In the features I have audited, the three biggest hidden multipliers are the same every time. Retries: a timeout followed by a retry doubles the cost of the slowest requests, which are also the largest. Agent loops: a task that takes eight steps sends the context eight times, so the cost grows roughly with the square of the conversation length, not linearly. And over-retrieval: teams pad the context with extra passages for safety, and pay for every one of them on every request, whether or not the model used it. None of these show up in a price-list estimate. All of them show up on the invoice. ## The five levers Once you see cost as a shape, the levers are obvious. Shorten what you send: trim the system prompt, summarize old turns, retrieve fewer and better chunks. Cache what repeats: a stable prefix can be cached by most providers, and a deterministic request can be cached by you. Route by difficulty: a small model handles the easy majority and a frontier model handles the hard minority. Bound the loops: a step budget is a cost budget. And cap output: ask for the shape you need, with structured outputs, instead of letting the model write an essay you will then parse. ## Cost is a latency and quality decision too The same shape that drives cost drives latency, because input tokens take time to process and output tokens take time to generate. Cutting a context in half usually makes the feature faster and cheaper at once. It sometimes makes it better, because a model given four relevant passages answers more precisely than one given twelve mixed ones. This is the rare case in engineering where the three axes point the same way, and it is worth taking advantage of before reaching for a smaller model. ## Budget per feature, not per month The practice I push on every team is to attach a cost budget to each feature at design time, expressed per user action: this summarization may cost this much, this agent run may cost that much. Then meter it in production with the same care as latency, per feature and per tenant, and alert when a feature drifts past its budget. The monthly invoice is a lagging indicator. The per-action cost is the leading one, and it is the one that tells you a prompt change or a loop bug is quietly burning money three weeks before finance notices. --- ## Auth is not a library decision 2025-12-04 · fullstack, security The first question I get on a new project is usually 'which auth library should we use?' It is the wrong first question. The library is the last decision, and by the time you make it, most of the important choices should already be written down. Authentication is a set of architectural commitments about identity, sessions, revocation and boundaries. A library implements those commitments. It cannot make them for you, and every auth incident I have investigated came from a commitment nobody made. ## Who owns the identity Decide whether your system is the source of truth for users or a consumer of someone else's. If you own it, you own password storage, reset flows, email verification and account recovery, and each of those is a place I have found critical vulnerabilities. If an identity provider owns it, you own the mapping from their subject id to your user record, and the failure mode becomes what happens when that mapping is wrong or when the provider is down. Neither is wrong. Both are decisions with consequences that outlast the library. A product that starts on a hosted provider and later needs to migrate off will discover that its user ids, its session format and its authorization rules were all shaped by that provider. ## Where the session lives There are two honest options. A stateful session stored in a database or a cache, referenced by an opaque cookie, can be revoked instantly and inspected at will, at the cost of a lookup on every request. A stateless token, signed and carried by the client, costs nothing to verify and cannot be revoked until it expires. Most teams choose stateless because it is easy, and then discover they cannot log a user out from the server, cannot invalidate sessions after a password change, and cannot see who is logged in. The usual compromise is short-lived access tokens with a refresh token stored server-side, so that revocation takes effect within minutes rather than never. That is fine, as long as you decide it rather than inherit it. The cookie itself has rules that do not depend on any library: HttpOnly so scripts cannot read it, Secure so it never travels in the clear, SameSite set deliberately, and a scope that matches your domains. I check these four attributes in the first ten minutes of any security review, and one is wrong more often than not. ## Where the check happens In Next.js, the temptation is to authenticate once in the middleware layer and trust that everything behind it is protected. That layer is good for cheap decisions: redirect an anonymous user, reject a request with no cookie at all. It is not the place for authorization, because it runs before the data is loaded and does not know what the user is trying to touch. The real check belongs next to the data. Every Server Component that renders private data, every Server Action that mutates it, every route handler, reads the session and verifies that this user may access this resource. Not just that they are logged in. That they own the order they are asking for. Insecure direct object references are the single most common vulnerability I find in modern React applications, and a login wall does nothing to stop them. I centralize that in a small data access layer: functions that take a session and an id, and either return the record or throw. Components never query the database directly. If the check is in one place, it is in every place. ## What the library is for Once identity ownership, session model, cookie policy and authorization placement are written down, the library choice is small. It handles the OAuth dances, the token formats, the provider adapters. Pick the one that matches the decisions you already made and that you can read the source of in an afternoon. If you pick the library first, it will make those decisions for you, quietly, and you will find out which ones it made during the incident. --- ## The hot partition problem and why sharding does not save you 2025-12-02 · system-design, data Sharding is sold as the answer to scale: split the data across N nodes, get N times the capacity. The sales pitch quietly assumes that traffic is spread evenly across keys. It almost never is. And when it is not, the node holding the hot key is your whole system, no matter how many others sit idle beside it. This is the hot partition problem. It is not exotic. I would say it is the single most common reason a "horizontally scalable" database falls over at a tenth of its theoretical capacity. ## How keys get hot Skew comes from a few recurring shapes. One big tenant: a SaaS with a thousand customers where one customer generates forty percent of the writes. If the shard key is tenant id, that tenant lives on one shard, and that shard is always the one on fire. One viral post: a social feature keyed by post id. Most posts get ten reads. One gets ten million in an hour. All ten million land on the same partition. Timestamps as keys: the classic mistake. If the partition key is the current hour, or an auto-incrementing id, every write in the system goes to the newest partition. You have paid for twenty nodes and you are using one, and tomorrow you will be using a different one. Notice that the problem in each case is a single key, or a small set of keys, that is hot by itself. Adding shards does nothing for that. A key lives in exactly one place. Doubling the cluster halves the load on the cold shards, which were fine, and leaves the hot shard exactly as hot as it was. ## Why the fix is never "more shards" People try it anyway, because it is the lever the vendor gave them. The result is a bigger bill and the same alert. The only ways out involve changing what a key means, changing where the read comes from, or changing when the write lands. ## The actual mitigations Salt the key. Instead of one partition for post 123, use post 123 with a suffix from 0 to 15, chosen at write time. Writes spread across sixteen partitions. Reads must query all sixteen and merge, so this trades read cost for write spread. It works well for counters, logs, and append-heavy data where reads are already aggregated. Split the hot entity. A tenant that is too big for one shard should not be one entity. Split by tenant plus region, tenant plus month, tenant plus workspace, whatever sub-boundary the domain already has. The split has to be something queries naturally include, or you have just moved the scatter-gather problem. Cache the hot read. The viral post does not need ten million database reads. It needs one read per second, held in memory, served ten million times. A hot key on the read side is a cache problem, and a cache with request coalescing solves it almost completely. Watch the key expire, though: the stampede after a hot key expires is the same storm you were avoiding. Aggregate writes behind a buffer. If the hot key is a counter, a like count, a view count, an inventory level for a flash sale, do not write every increment to the row. Accumulate in memory or in a queue, flush every second. One thousand increments become one write. You lose a second of freshness and gain three orders of magnitude of headroom. For inventory this needs care, because over-selling is a real cost, so pair the aggregate with a hard reserve step for the last units. Separate the whale. Sometimes the honest answer is that one tenant is a different workload. Give them their own database, their own cluster, their own cost line. Multi-tenancy is a sharing decision, and the biggest customer is often the one you should not share with. ## What to look at this week Pull the per-partition metrics from whatever you run. If you cannot get them, that is the first finding. If you can, sort by writes. If the top partition does more than a few times the median, you already have a hot partition. You have just not had the traffic day that turns it into an incident. --- ## The year I lost every contract 2025-12-01 · career, life In 2020 I had a small technology company in Toledo, a set of contracts that paid the bills, and a wife who was pregnant. Within a few weeks of the pandemic starting, every contract was gone. Not reduced. Gone. I remember the day the last one ended. I usually cycled home. That day I walked, because I needed the time to cry before I got to the door and had to tell Natália. I want to write about what came after, not because the story is unusual, many people lived a version of it that year, but because what I rebuilt was different from what I lost, and the difference is the useful part. ## What I had was not a business, it was a list of clients The first honest thing I had to admit was that I did not lose a business. I lost a list. There was nothing under the contracts: no process that made me replaceable, no written way of working, no asset a new client could evaluate without talking to me first. When the clients stopped, there was nothing left to sell. That is a System Design failure, and I recognised it as one because it is exactly what I find when I audit companies. Everything depends on one node. The node goes down and there is no fallback. I had designed my own livelihood with a single point of failure and had not noticed, because the node was me. ## Rebuilding with a different shape I did not rebuild the same thing faster. I rebuilt something with a different shape, and the shape had three properties. - Evidence before relationship. Every piece of work had to leave an artefact that could be evaluated without me: a report, a design document, a review someone could forward. - Depth over breadth. Instead of doing everything for a few clients, I did a narrow thing, finding failures in systems, for many. Depth compounds. Breadth just fills a calendar. - Written down, always. Decisions, reasoning, what was tried and what happened. This is what makes the work outlive the contract. That shape became fgxdev.com. It is not a big studio. It is a small one built so that losing any one client is a bad month, not a catastrophe. ## Faith, and what it did and did not do My Christian faith became important to me during that year. I want to be precise about this because it is easy to be sentimental. Faith did not bring the contracts back. It did not tell me which clients to call. What it did was change the question I was asking from "why me" to "what now", and it gave me a reason to get up and do the boring next step on days when there was no visible reason to. That is all, and it was enough. ## What I would tell someone in the middle of it If you are reading this in the week you lost everything, here is the only thing I can offer that is concrete: audit what you lost the way you would audit a system. Draw it. Find the single points of failure. Notice that the thing you are grieving was probably more fragile than you thought, and that the version you build next can be less fragile on purpose. Then take the walk. Cry if you need to. Tell the person waiting at home the truth. And the next morning, write down the first small thing you can produce that does not depend on anyone calling you back. That is the first node of the new system. Everything I have built since, including the research project I am working on now, started with that one written page. --- ## Fine-tuning, retrieval or prompting: a decision I make weekly 2025-11-29 · ai Almost every AI feature proposal that lands on my desk arrives with a preferred technique already attached. Someone wants to fine-tune because it sounds serious. Someone wants RAG because they read that it fixes hallucination. Someone wants a bigger prompt because it is what they know. The technique is rarely chosen from the problem. So I have a short decision procedure that I run every week, and it starts with a single question: what does the model need that it does not already have, and how often does that thing change? ## The decision tree A model can be missing three different things. It can be missing instructions: it does not know what you want, in what shape, with what tone. It can be missing knowledge: facts about your domain, your data, your customers, this week's policy. Or it can be missing behavior: a consistent way of acting that instructions describe poorly, like a house style, a classification boundary that is hard to put into words, or a format so specific that examples are cheaper than rules. Instructions are fixed with prompting. Knowledge is fixed with retrieval. Behavior is fixed with fine-tuning. Most problems are the first two, and most fine-tuning proposals I see are attempts to fix the first two with the third. ## Prompting first Prompting is where I always start, because it is the cheapest to change and the fastest to evaluate. A clear task description, a schema for the output, a handful of well-chosen examples and an explicit statement of what to do when unsure will get a frontier model most of the way on most tasks. The mistake is to stop iterating too early or too late. Too early means declaring the model incapable after two attempts. Too late means a four-thousand-token prompt that has become an unmaintainable rulebook, which is the signal that the missing thing was never instructions. ## Retrieval when facts change If the model needs to know something specific and that something changes, retrieval is the only sane option. Product catalogs, internal documents, customer histories, regulations, anything with a date on it. Fine-tuning facts into a model is expensive, slow, and stale the moment the facts change, and it does not give you a citation. Retrieval gives you a citation, updates when the source updates, and can be permissioned per user. The cost is that you now own a search system, which is real engineering. But that engineering is well understood, and it fails in ways you can see. ## Fine-tuning when behavior needs shaping I reach for fine-tuning when I have a large, clean set of examples of the exact behavior I want, when prompting has plateaued with the eval set proving it, and when the task is stable enough that the investment pays back. Typical cases: a small model that must match a frontier model's output quality on one narrow task, at a fraction of the cost and latency. A format or style that examples convey better than words. A classifier on domain-specific inputs where a thousand labels beat any description. What fine-tuning does not do is add knowledge reliably or fix a task that was never specified clearly. The eval set decides, not the enthusiasm. ## The weekly ritual In practice the decision is not once per feature. It is revisited as the feature matures. A feature starts with prompting, because that is how you learn what the task even is. It grows a retrieval layer when the eval set shows failures caused by missing facts. It might get a fine-tuned small model for cost once the behavior is stable and the examples have piled up. And every provider release resets some of this, because a better base model can make yesterday's fine-tune unnecessary. That is fine. The technique was never the point. The eval set is, and the technique is whatever makes the numbers go up this week. --- ## A post-mortem template that produces change 2025-11-27 · infra, reliability I have read hundreds of post-mortems, mine and other people's. The good ones share a property that has nothing to do with writing quality: thirty days later, something in the system is different because of them. The bad ones are often better written. They have a timeline, a root cause, a list of action items, and none of the action items ever ship. The difference is structural. A post-mortem that produces change is built around questions that force uncomfortable specifics. Here is the template I use, and why each section exists. ## Impact, in the customer's units Not 'the API was degraded for 42 minutes'. Instead: how many users were affected, what they could not do, how many orders or appointments or messages were lost or delayed, and what it cost, in money or trust or support hours. If the number is not known, write 'unknown' and add finding it to the actions. The point is to make the incident's weight visible to the people who decide what engineering works on next. ## Timeline, with the gaps named A timeline from first cause to full recovery, with timestamps. The useful part is not the events but the gaps between them. How long between the failure starting and someone noticing? Between noticing and understanding? Between understanding and acting? Between acting and recovering? Each gap is a separate problem with a separate fix. A forty-minute gap between failure and detection is an alerting problem. A forty-minute gap between detection and understanding is an observability problem. A forty-minute gap between understanding and recovery is a deploy or runbook problem. ## Contributing causes, plural There is almost never one root cause. There is a trigger, and there are three or four conditions that turned the trigger into an outage. The bad deploy was the trigger. The missing health check let it reach all instances. The absent canary meant it reached them at once. The alert threshold was too high to fire. The runbook pointed at an old dashboard. Every one of those is a cause, and fixing only the trigger guarantees the next trigger will find the same conditions waiting. I ask 'what would have made this a non-event' for each layer: prevention, detection, mitigation, recovery. Each layer usually yields one concrete change. ## Actions with an owner, a date and a test This is where most templates fail. 'Improve monitoring' is not an action. 'Add an alert on p99 latency above 2 seconds for 5 minutes on the checkout endpoint, owned by a named person, done by a date, verified by triggering it in staging' is an action. Every item needs all four parts: what exactly, who, when, and how we will know it worked. Then the hard rule: actions go into the same tracker as product work, with the same priority process. If they live in a separate document, they die in that document. If a decision is made not to do one, that decision is written down with the name of whoever made it, so that the next post-mortem can reference it honestly. ## The questions that make it honest - What did we believe about the system that turned out to be false? - What signal was available before the incident that nobody was looking at? - What did the person on call have to improvise that should have been written down? - Which of these actions would we have rejected as unnecessary a month ago, and why? - What is the next incident this same set of conditions would produce, if we fix nothing? That last question is the one I insist on. It turns the post-mortem from a report about the past into a prediction about the future, and predictions are what get budget. One last thing, about blame. The person who pushed the bad deploy is not the cause. The system that let one push become an outage is the cause. Say that clearly, and people will tell you what actually happened instead of what protects them. But blameless does not mean nothing changes. It means the changes land on the system, not on the person, and the template above is how you make sure they land at all. --- ## The monolith is not the problem, the coupling is 2025-11-25 · architecture I have lost count of the number of times a team told me 'we need to break up the monolith' when what they actually had was a coupling problem that would follow them into any topology they chose. The monolith is a deployment unit. Coupling is a design property. You can have a monolith with clean internal boundaries that deploys in four minutes and lets ten teams work in parallel. You can have forty microservices that all have to deploy together because they share a database and a set of implicit assumptions. The second one is a monolith too. It is just a monolith with network latency and no transactions. ## What coupling actually is Coupling is the degree to which a change in one place forces a change somewhere else. That is the whole definition. It has nothing to do with process boundaries. The kinds that hurt most, in my experience, are these. Data coupling: two modules read and write the same tables, so a schema change in one is a bug in the other. Temporal coupling: A must run before B, and nobody wrote that down, so it works until someone reorders a job. Semantic coupling: two modules both 'know' that status 3 means 'shipped', encoded as magic numbers in each. Build coupling: a change to a utility package forces a rebuild and retest of everything that imports it, which is everything. None of these are solved by putting an HTTP call between the modules. Data coupling becomes a shared database across services, which is worse. Temporal coupling becomes a race condition across the network. Semantic coupling becomes two services that disagree about what status 3 means, and now the disagreement is in production logs instead of a compiler error. ## The honest diagnostic Before I recommend any split, I ask for three things. First, the dependency graph between modules, not between services. Most teams cannot produce it. That alone tells me the coupling is undocumented, which means it is uncontrolled. Second, the last twenty pull requests. I count how many files each one touched and in how many top-level folders. If a typical change to 'orders' touches 'orders', 'billing', 'notifications' and 'shared', the boundaries are not where the folders are, and a service split along the folders will require the same four changes across four repos with four deploys. Third, the database schema with foreign keys drawn. If the orders table joins to users, products, addresses, payments and a shared audit table, and every module queries all of them, there is no seam to cut along. Splitting here would mean either duplicating data with sync jobs or making the services call each other synchronously for every read. Both are worse than what you have. ## Fix the coupling inside the monolith first The work that actually helps is boring and can be done without changing the deployment model at all. Give each module its own set of tables and forbid cross-module queries at the code review level, then at the lint level, then at the database permissions level. Replace direct function calls across modules with an explicit interface that lives in one file per module. Move shared enums and status codes into the module that owns the concept, and make everyone else import them from there. Do this for six months. Measure it. Count the files per PR again. When 'orders' changes stop touching 'billing', you have earned the right to ask whether they should be separate services. And most of the time, once the coupling is gone, the answer is 'not yet'. The monolith with clean boundaries deploys fast, tests in one process and has real transactions. Those are enormous advantages that teams throw away for a problem they have not solved. ## When the split is actually justified There are honest reasons to extract a service. A module needs a different runtime, a different language, a different scaling profile, a different compliance boundary, or a different release cadence that is blocking others. Those reasons are specific and measurable. 'It is too big' is not one of them. 'It is hard to change' usually means coupling, and coupling travels. Fix the coupling. Then look at the monolith again. You may find it is not the problem you thought it was. --- ## Six hundred projects later: what repeats 2025-11-22 · audit, career Somewhere past the five hundredth project I stopped being surprised. Not because systems stopped failing, but because they kept failing in the same handful of ways, regardless of language, framework, team size or industry. A marketplace, a clinic scheduling app, a logistics platform and a fintech backend share almost nothing on the surface and almost everything underneath. Here is the list. It is short because the truth is short. ## The list - No timeouts on outbound calls, so one slow dependency takes the whole system down with it - Unbounded queries, a list endpoint with no limit that works at a thousand rows and dies at a million - Writes that are not idempotent, so every retry, double-click or replayed webhook creates a duplicate - Errors caught and swallowed, so the system keeps going in a state nobody understands - Secrets in configuration files, in repositories, in logs, in error messages sent to the client - Backups that exist and have never been restored, which means they are a hypothesis - One person who knows how the deploy works, and the deploy only works when they are awake ## Why these and not others Each of these is invisible while the system is small. A missing timeout does not matter when the dependency is fast. An unbounded query is fine at launch. A non-idempotent write only duplicates when something retries, and nothing retries until the first outage. That is the pattern: these failures are dormant by construction, and the moment they wake up is the moment the business finally has enough traffic to matter. The success of the product is the trigger for the failure. Sophisticated bugs, the ones in clever algorithms or subtle concurrency, are rare by comparison and usually caught by the people who wrote the clever code, because they were paying attention. The boring failures survive because nobody was paying attention to the boring parts. ## What I stopped believing I stopped believing that framework choice matters much. I have seen the same seven failures in Rails, Django, Spring, Express and Next.js. I stopped believing that more tests fix it, because tests exercise the paths people thought of, and these are the paths nobody thought of. I stopped believing that a senior team is immune. Senior teams make these mistakes under deadline pressure exactly like junior ones, they just feel worse about it afterwards. ## What I started doing I check the list. Every project, every time, before I read anything else. It feels mechanical and slightly insulting to the team, and it finds something in the large majority of systems I look at. The check takes half a day. The failures it prevents take weeks and happen at the worst possible moment. That trade is so lopsided that I no longer understand why it is not universal. I also started building the countermeasures in from day one on my own projects: a default timeout on every HTTP client, a maximum page size enforced at the query layer, an idempotency key on every mutating endpoint, a restore drill on the calendar, a deploy that a new hire can run from a document. None of it is impressive. All of it is the difference between a system that survives its own success and one that does not. ## The honest conclusion Six hundred projects taught me that the frontier of software failure is not where the industry looks. It is not in the new framework or the model or the architecture pattern. It is in the same seven places it was a decade ago, waiting for the traffic to arrive. The most valuable thing I can do for a client is often the least glamorous: walk the list, find the sleeping failure, and wake it up in a report instead of in production. --- ## Feature flags as architecture 2025-11-20 · fullstack, engineering A feature flag is a conditional that lives in production and can change without a deploy. That is a powerful thing to have, and every team I have worked with was glad they added flags. It is also a second control flow, invisible in the code, that decides what your system does. Most teams treat it as a convenience. It is architecture. The difference shows up in the audit. A codebase with forty flags, half of them permanently on, three of them referenced in code that no longer exists, and one that turns off payment retries and has been on since an incident two years ago. Nobody knows which combination of flags production is running. The bug that only happens for some users is not a race condition. It is a flag. ## Flags have types The first discipline is to admit that not all flags are the same thing. A release flag hides unfinished work and should live for weeks. An experiment flag splits traffic and should live until the experiment concludes. An operational flag is a kill switch for a dependency and should live forever. A permission flag gates a feature per customer and is really a product entitlement, not a flag at all. Mixing them is where the trouble starts. A release flag that is never removed becomes an accidental operational flag. An entitlement modeled as a boolean flag becomes impossible to bill for. I name flags with their type as a prefix and I give each type a different lifecycle rule. ## Evaluate once, at the edge of the request The most common implementation mistake is checking the flag service everywhere. Deep in a domain function, in a React component, in a queue worker. Each check is a network call or a cache read, and each one can see a different value if the flag changes mid-request. Half a checkout runs with the new pricing and half with the old. The rule I enforce: flags are evaluated once, at the entry point, into a plain object of booleans and variants that travels down with the request. In Next.js that means reading the flags in the layout or the Server Action, once, and passing the result as an argument or through context. The domain code receives a decision, not a flag client. It becomes testable without mocking a service, and the request is consistent from start to finish. For Server Components, that object is also what you serialize to the client parts that need it. The client never calls the flag service. It receives the decisions the server already made, which also means the client bundle does not contain a flag SDK and its secrets. ## Flags change the cache key A page rendered with flag A on and cached is a different page from the same route with flag A off. If the cache key does not include the flag state, a user in the experiment sees the control page from the cache, and the experiment results are noise. This is the part everyone forgets. Every flag that changes what a route renders is part of that route's cache identity, on the framework's cache and on the CDN. Either the flag state is in the key or the route is not cached. There is no third option, and I have seen experiments run for a month whose data was meaningless because of this. ## Removal is part of the feature A flag is not done when the feature ships. It is done when the flag is deleted and the losing branch is gone. I put the removal ticket in the same pull request that adds the flag, with a date, and I fail CI on any release flag older than ninety days. That sounds harsh. It is what keeps the count at ten instead of forty. The kill switches stay, but they are tested. Once a quarter, flip each operational flag in staging and confirm the system degrades the way the runbook says. A kill switch nobody has tried since the incident that created it is a hypothesis, not a control. Flags are worth it. They are also a second program running alongside your first one, and it deserves a design. --- ## The queue you did not know you had 2025-11-18 · system-design Every system has more queues than its architecture diagram shows. The diagram shows the message broker. It does not show the connection pool, the thread pool, the socket accept backlog, the database lock wait, the load balancer buffer, or the retry loop in a client. Each of those is a queue, and each one has a size, a wait time, and a failure mode you never chose. ## Where the hidden queues live Start at the edge. The operating system holds incoming connections in an accept backlog before your server touches them. The web server has a worker pool; requests beyond it wait. Your application has a database connection pool; a request that needs a connection when all are busy waits, invisibly, inside a library call. Inside the database, writes to the same row queue on a lock. A long transaction on an orders table turns every other write to those rows into a line of waiting transactions, and the app sees it as "the database is slow". Outbound, an HTTP client has a connection limit per host. A downstream service that gets slow keeps those connections busy, and the next call queues in the client. A logging library with an async appender has a buffer. A serverless platform has a concurrency limit and queues invocations past it. None of these appear in a diagram, and all of them obey the same law. ## The law Little's law: the number of items in a system equals the arrival rate multiplied by the time each item spends there. It sounds academic and it is the most practical formula in operations. Take a connection pool of 20 connections. If queries take 10 milliseconds, the pool sustains about 2,000 queries per second. If a slow query pushes the average to 100 milliseconds, the same pool sustains 200. Nothing changed in the traffic; the pool just became ten times smaller in effect. The requests that do not fit wait, their wait adds to the next request's time, and the queue grows until something times out. This is why "the database is slow" and "the app is down" arrive together. It is not two incidents. It is one hidden queue filling. ## Finding them When I audit a system, I ask for every place a request can wait, and I ask three questions about each: how big is it, how long can something wait in it, and what happens when it is full. Most teams can answer for their broker and for nothing else. Then I check whether the queue is measured. A pool exposes active, idle and waiting counts. A web server exposes queued requests. The database exposes lock waits. If those numbers are not on a dashboard, the queue is invisible, and an invisible queue is one you discover during an incident. ## Making them explicit The fix is rarely to remove the queue. It is to give it a size and a timeout that you chose, and a metric that you watch. Connection pools get a maximum wait time, after which the request fails fast. Thread pools get a bounded task queue. HTTP clients get per-host limits that match what the downstream can serve. Long transactions get split so lock waits stay short. Serverless functions get reserved concurrency so one hot endpoint cannot starve the rest. Once the hidden queues have limits, the system fails in the place you chose, with an error you recognise, instead of in a random place with a timeout. ## The mental shift Stop thinking of latency as a property of code and start thinking of it as time spent in queues. A request that "takes 800 milliseconds" usually does 80 milliseconds of work and 720 milliseconds of waiting in places nobody drew. Find the waiting, and you find both the bottleneck and the failure mode. The queue you did not know you had is the one that will page you. --- ## Reproducibility is an engineering problem 2025-11-15 · science, engineering Engineers solved 'it works on my machine' years ago, not with discipline, but with tooling that made the disciplined path the default. Lock files, containers, immutable builds, CI that runs on a clean machine. When I hear a researcher say an analysis cannot be rerun because the student who wrote it graduated, I do not hear a moral failure. I hear a missing build system. ## What reproducible actually means Same inputs, same code, same result, on a different machine, run by a different person, a year later. Each clause is a separate engineering requirement. 'Same inputs' needs the raw data to be immutable and addressable by a hash, not a filename. 'Same code' needs the exact versions of every library, not 'pandas' but 'pandas 2.2.1 with numpy 1.26.4'. 'Different machine' needs the environment declared, not assumed. 'A year later' means the dependencies still have to resolve, which means you pin or vendor them, because the internet forgets. Most analyses fail on the second clause. A random seed that was set interactively and not in the script. A library update that changed a default. A path that only exists on one laptop. ## Raw data is read-only, forever The first rule I give any team, in software or in science, is that raw data is never edited. It lands in a directory that nobody can write to, with a checksum recorded next to it, and every cleaning step is a script that reads from there and writes somewhere else. If a value is wrong in the raw data, you document it and correct it in a transformation, so the correction is visible and reversible. This single rule eliminates the most common failure I see: someone fixed a typo by hand in the spreadsheet, the fix was right, and now no one can prove what the original looked like or whether other cells were touched in the same session. ## The analysis is software, treat it that way An analysis notebook is a program. Programs need version control, a way to run them from the command line with no human clicking cells in order, and at least one test that asserts something known about the output. The test does not have to be sophisticated. 'The cohort has 412 rows after filtering' is a test. 'The mean age is between 40 and 80' is a test. When the number changes, you find out the day it changes, not at peer review. Then put the whole thing in a container and run it in continuous integration on every commit. If it produces the same figures on a clean machine, you have reproducibility. If it does not, you have found the hidden dependency now, cheaply, instead of later, expensively. ## Determinism is a choice you make early Some computation is non-deterministic by nature: parallel reductions, GPU kernels, sampling. You cannot always make it bit-for-bit identical, but you can decide what 'the same result' means and check it: tolerance bands, fixed seeds, statistical equivalence. What you cannot do is discover this after the fact. I have audited pipelines where two runs disagreed in the third decimal place and nobody could say whether that was noise or a bug, because nobody had decided in advance. ## Why an engineer cares Reproducibility is not a virtue to be exhorted. It is a property of a system, and systems get properties through design. Every practice above is standard in a decent software team and requires no new science. What it requires is someone treating the lab's computational work as infrastructure and giving it the same care we give a payments service. That is a role I recognize, and a job I think is worth doing. --- ## Observability is a product feature 2025-11-13 · infra, observability A customer writes in: 'my order disappeared'. What happens next tells me more about a system than any architecture diagram. In a well-built system, someone types the order id into a search box and, thirty seconds later, has the full story: created at 14:02, payment authorized, inventory reservation timed out, compensating cancellation fired, email queued but bounced. In most systems I audit, what happens next is a two-hour investigation across three dashboards and a database console, ending in 'we think it was the payment provider'. The difference is not tooling. Both teams have logs and metrics. The difference is that the first team treated observability as something the product does, not something the platform provides. ## The question you are building for Observability has one job: answering questions you did not know you would ask. Metrics answer the questions you predicted. Alerts answer the questions you feared. But the customer support ticket, the fraud investigation, the 'why is this tenant slow' question, those are unpredictable, and only rich, correlated, searchable context answers them. So the design target is simple: given any business identifier (order id, user id, tenant id, request id), how long does it take to reconstruct what the system did? If the answer is measured in minutes, you have observability. If it is measured in meetings, you have logs. ## Three things that make it work The first is a correlation id that survives every boundary. It is generated at the edge, attached to every log line, every queue message, every outbound HTTP call, every database comment if you can afford it. Without it, you have events. With it, you have a story. In a Next.js application this means the id is created in middleware, passed through server actions and API routes, and injected into the logger context, not passed by hand. The second is that business identifiers are first-class fields, not text inside a message. A log line saying 'processing order 8812' is nearly useless. A log line with order_id as a structured field, alongside tenant_id and user_id, is searchable, filterable and joinable. The whole value is in the fields. The third is that the important events are emitted on purpose. Not 'entered function', but 'payment authorized', 'reservation failed', 'refund issued'. These are domain events. They are the same events your product team would list if you asked them what matters. Emitting them as structured logs or spans gives you a business-level timeline for free. ## Why it belongs on the roadmap When observability is treated as infrastructure, it competes with nothing and gets nothing. When it is treated as a feature, it competes with other features, and it wins more often than people expect. Consider what it delivers: support tickets resolved in minutes instead of hours, incidents whose root cause is found on the first day, fraud patterns visible before they become losses, performance regressions caught per tenant instead of averaged away. In one marketplace I reviewed, the single most valuable engineering investment of the year was an internal page that took an order id and showed its timeline. It was not glamorous. It removed an entire category of escalation. ## The minimum I recommend Structured JSON logs with a stable set of fields. A correlation id from the edge to the database. Domain events emitted explicitly at every state transition that matters to the business. Retention long enough to investigate last month, not just last night. And one internal tool, however ugly, that turns a business id into a timeline. That is a few days of work on a small system and it pays back in the first serious incident. The product team will not ask for it. You should build it anyway, and then show them what it can do. After that, they will never let you remove it. --- ## API versioning is a relationship, not a number 2025-11-11 · architecture, engineering Every API versioning debate I have sat through started in the wrong place. Path or header? v1 or a date? Semantic or incremental? Those are formatting questions. They decide how a version is spelled, not what it means, and a team that answers them first usually ships a v2 that breaks the same consumers v1 did. A version is a promise made to the people who built on top of you. The number is just the label on the promise. If you do not know who those people are, what they call, and how fast they can change, the label is decoration. ## Who is on the other side The first thing I ask when auditing an API is not how it is versioned but who consumes it. The answers fall into three groups, and each needs a different kind of relationship. Internal teams can be talked to. You can open a pull request in their repo, or at least a ticket with a deadline. Breaking changes are negotiable, and a version bump is often more ceremony than protection. Known external partners have contracts, sometimes literal ones. They deploy on their own schedule, usually slower than yours. A breaking change here costs a meeting, a migration window and goodwill. Anonymous consumers, behind public API keys, mobile apps in the wild, or a widget someone embedded three years ago, cannot be talked to at all. For them the version is the only channel you have, and you must assume the old one will be called forever. Most APIs serve all three. Versioning them with a single global scheme treats the team next door like an anonymous mobile app from 2022. That is either too rigid for the first or too loose for the third. ## What actually breaks people Consumers do not break on version numbers. They break on shapes. In order of frequency: a field removed, a field renamed, a nullable field that used to always be present, an enum with a new value the client's switch statement did not expect, a validation that got stricter, and a default that changed. Most of those are not 'breaking' by the strict definition many teams use. And yet each of them takes down a client that trusted the old behavior. A versioning policy that only covers removed fields protects against the least common failure. So my rule is: expansion is free, contraction is a version. You can add fields, endpoints, optional parameters and enum values, as long as you have documented that clients must tolerate unknown values. You cannot remove, rename, tighten or change meaning without a new version and a migration path. And 'meaning' includes semantics like whether an amount is in cents or in units. ## The mechanics I actually recommend The mechanics almost pick themselves. Put the major version in the URL path, because it is visible in logs, in curl commands and in the browser, and because anonymous consumers can find it without reading headers. Do not version minor changes at all; make them additive. Keep at most two majors alive, and give the old one an end date the day the new one ships. Not 'deprecated', which nobody reads, but a date, in the docs and in a response header, with a warning that grows louder as it approaches. Then instrument it. A version you cannot measure is a version you cannot retire. Tag every request with its version and consumer identity, so that when the end date comes you know exactly which three partners are still on v1 and can call them by name. ## The relationship part Here is what the number cannot do. It cannot tell a partner that a change is coming, or why. It cannot give them a sandbox with the new shape six weeks early. It cannot tell you that one of your biggest consumers is a batch job that runs once a quarter and will not notice anything until it fails in production. Those things are done by people, with changelogs, a deprecation page and, for the consumers that matter, a conversation. The teams that version APIs well treat each major as a project with stakeholders, not as a branch. They know their consumers' deploy cadence. They ship a migration guide before the code. They watch the traffic move and follow up with whoever has not. The teams that version badly have beautiful v3 URLs and a v1 that will never be turned off, because nobody knows who still calls it, and turning it off feels like cutting a wire in the dark. An API version is a relationship with the people who trusted you enough to build on your work. Spell it however you like. Just do not confuse the spelling with the promise. --- ## RAG is a data pipeline with a language model at the end 2025-11-08 · ai, system-design When a team tells me their RAG feature gives bad answers, I ask to see the ingestion job before I ask to see the prompt. Nine times out of ten the problem is there. The PDF parser dropped every table. The chunker split a clause from the sentence that negated it. The index has three copies of last year's policy and none of this year's. The model at the end did exactly what it was told with exactly what it was given. It was given garbage. ## Where RAG features actually fail Retrieval-augmented generation is a pipeline: acquire documents, parse them, split them, embed them, store them, query them, rank them, assemble a context, generate. Only the last step involves anything that could reasonably be called AI, and it is the step with the fewest bugs, because the model is a well-tested component you did not write. Every other step is code you wrote, in a hurry, against data you did not look at closely. That is where I look first, and it is where the failures nobody else found are usually hiding. ## Ingestion is the product A document pipeline I audited had a beautiful prompt and a re-ranker and a hybrid search layer, and it could not answer questions about pricing, because pricing lived in tables and the parser flattened tables into word soup. No prompt fixes that. The teams with the best RAG features are boring about ingestion: they keep the source, they version the parser, they store metadata like section, date and owner alongside every chunk, they detect duplicates, and they re-index when the source changes. They treat it like an ETL job, with the same tests and the same alerts, because that is what it is. ## Retrieval is a search problem, and search is old Vector similarity is one signal. It is a good signal for meaning and a poor one for exact identifiers, dates, negations and rare terms. Good retrieval combines it with the things search engineers have known for decades: lexical matching, metadata filters, recency, source authority and a ranking step that is evaluated against real queries. The question to ask is not whether the embedding model is good. It is whether, for the fifty questions your users actually ask, the right passage lands in the top few results. Measure that directly, with a labeled set, before touching anything downstream. ## The model is the last mile Once the right passages arrive, the generation step has a narrow job: answer from the context, cite what it used, and say so when the context does not contain the answer. That last behavior is where the system prompt matters, and it is testable: feed it a question the corpus cannot answer and check that it declines. The model should be small in your mental model of the feature. It is a formatter with judgment, sitting on top of a search system that carries the actual weight. ## What I check first in a RAG audit I sample twenty chunks from the index at random and read them. Not the prompt, not the dashboards, the chunks. If they are coherent, self-contained and traceable to a source, the pipeline is probably healthy and I move on to retrieval quality. If they are fragments with no context, split mid-sentence, missing the heading that gives them meaning, I stop the audit right there, because nothing downstream can be trusted until the data is. That is a data engineering finding, and the fix is data engineering. The language model was never the problem. --- ## Caching in Next.js without surprises 2025-11-06 · fullstack, nextjs Every Next.js incident I have been called into that involved stale data had the same root cause. Nobody on the team could say, for a given page, which cache layers it passed through and who was responsible for clearing each one. That is the whole problem. Not the framework, not the defaults. A cache you cannot describe is a cache you cannot debug. ## Know the four layers There are four places a value can be cached in a Next.js application, and they have different lifetimes and different owners. Request memoization dedupes identical fetches during a single render, it lives for one request and you never have to invalidate it. The data cache stores the results of fetches or cached functions across requests and across deploys until you revalidate them. The full route cache stores the rendered HTML and payload of static routes at build time. The client router cache keeps visited route segments in the browser's memory during a session. Since Next.js 15, the defaults are conservative: fetch is not cached unless you ask, GET route handlers are not cached, and the router cache treats dynamic pages as stale immediately. That was the right change. It means that in a modern project, a stale value is something you explicitly opted into, and you can find the line that did it. Then there is the fifth layer everyone forgets: the CDN in front of the application, driven by Cache-Control headers. The framework does not manage it and it will happily serve a page the framework thinks it revalidated. ## Cache by name, not by hope My rule is that nothing enters the data cache without a tag. The `'use cache'` directive lets you cache a function or a whole component, and `cacheTag` lets you name what it depends on. A product page gets the tag for that product and the tag for its category list. An order summary gets the tag for that order. Tags are what make invalidation possible. When a Server Action updates a product, it calls `revalidateTag` with the product tag and the category tag, and every cached fragment that depends on those is marked stale. No cron job, no guessing which paths to purge, no clearing the whole cache because you were not sure. I keep a small module that builds tag strings from ids, so that the tag for product 42 is spelled the same way in the read and in the write. A typo in a tag string is an invalidation that silently never happens, and I have found that exact bug in production more than once. Time-based expiry is the fallback, not the strategy. `cacheLife` profiles are fine for data you do not own, such as an exchange rate from a third party, where you cannot know when it changed. For your own data, you know exactly when it changed, because you changed it. Invalidate at the mutation. ## The surprises, catalogued The most common surprise is per-user data in a shared cache. A cached function that reads the session inside its body will serve the first user's data to everyone. Anything cached must take the user id as an argument so that it becomes part of the cache key, or it must not be cached at all. The second is the accidental static page. A page that does not read cookies, headers or search params and has no dynamic data will be prerendered at build time and never change until the next deploy. That is wonderful for a landing page and disastrous for a pricing page that reads from a database through a cached function nobody tagged. The third is read-your-own-writes. A user submits a form, the action revalidates the tag, but the user is redirected to a page that was served from the CDN with a sixty-second max-age. The framework did its job. The header did not. Set Cache-Control on purpose, per route, and keep it short or absent for anything a user can change. ## The checklist I run Before a launch, I open every route and write one line for each: dynamic or static, which tags it depends on, which mutations revalidate those tags, and what Cache-Control it sends. If I cannot fill in a line, that route is going to surprise someone. It takes an hour. The alternative is the ticket that says 'the data is wrong sometimes', and that ticket takes a week. --- ## Exactly-once is a promise nobody keeps 2025-11-04 · system-design, engineering Every few months a broker, a stream processor or a managed queue announces exactly-once delivery. Every time, the fine print says something narrower. I have stopped reading the headline and started reading the fine print, because the headline is a promise and the fine print is where the engineering lives. ## Delivery is not processing There are two different promises hiding behind the same phrase. Exactly-once delivery means the message arrives at the consumer once. Exactly-once processing means the effect of the message happens once: the row is inserted once, the email is sent once, the balance is debited once. Vendors sell the first. You need the second. And the first does not give you the second, because your consumer can crash after doing the work and before telling the broker it is done. ## The two generals, without the math Two generals on opposite hills need to attack at the same time. They can only communicate by messenger, and messengers get caught. General A sends 'attack at dawn'. Did it arrive? A does not know until B confirms. Did B's confirmation arrive? B does not know until A confirms the confirmation. There is no finite number of messages after which both are certain. That is the whole problem, and it applies to a consumer telling a broker 'I handled message 4712'. The acknowledgement can be lost. The broker must then either redeliver (at-least-once) or hope for the best (at-most-once). There is no third option over a lossy network. Exactly-once is something you construct on top, not something the network gives you. The acknowledgement is the hard part, always. ## What effectively-once actually looks like The pattern that works is boring: at-least-once delivery plus an idempotent consumer. The broker may redeliver. The consumer makes redelivery harmless. Together they produce effectively-once, which is the honest name for the guarantee people want. Idempotent means running the same message twice leaves the system in the same state as running it once. Some operations are naturally idempotent: set status to shipped, upsert a row by primary key. Most are not: increment a counter, append a line, send a notification, charge a card. For those you need a dedup key. A dedup key is a stable identifier for the intent, not for the delivery attempt. If the producer generates a new UUID per retry, you have a key for nothing. The order id plus the event type is a good key. The message id assigned by the broker is often not, because some brokers assign a fresh one on redelivery. Store the key in the same database transaction as the side effect. Insert into processed_events and update the balance in one transaction. If the insert fails on the unique constraint, the work was already done, so acknowledge and move on. If the transaction fails, nothing happened, so let the broker retry. This is the one place where 'same transaction' is not optional. ## The retention window is a design decision You cannot keep dedup keys forever, and this is where most implementations get sloppy. The window must be longer than the longest possible redelivery. That includes the broker's retry policy, dead letter replays, and the human who reprocesses last week's messages during an incident. Seven days is a common starting point; I have seen 24 hours be wrong the first time someone replayed a partition. Pick the number deliberately and write it down next to the retry configuration, because they are one decision, not two. ## Transactional producers cover one leg Kafka-style transactional and idempotent producers are real and useful. They make sure a producer retry does not write the same record twice into the log, and that a consume-transform-produce loop commits offsets and outputs atomically. That is exactly-once inside the log. It says nothing about the moment your consumer calls a payment API, writes to Postgres or sends an email. The moment the effect leaves the broker's transaction boundary, you are back to at-least-once plus idempotency. Most real effects live outside that boundary. So my rule is simple. Assume every message will arrive at least twice, and sometimes out of order. Make the second arrival a no-op by construction. Put the dedup key and the effect in one transaction. Choose a retention window longer than any replay you will ever run. Only then read the vendor's exactly-once page, as a nice optimization of the leg it actually covers. When a team tells me their system is exactly-once, I ask one question: what happens if the consumer dies between the database commit and the ack? If the answer is a shrug, the answer is 'twice'. --- ## Strict mode is the cheapest code review you will ever get 2025-10-30 · fullstack, typescript I have read a lot of code review guidelines. Most of them ask the reviewer to check for things a compiler could check in a millisecond: a value that might be null, an argument in the wrong position, a case nobody handled. Humans are bad at this. They get tired, they skim, they trust the author. The compiler does not. Turning on strict mode in TypeScript is hiring a reviewer who reads every line of every file on every keystroke, never skips a case and does not care that the deadline is tomorrow. The price is one line in tsconfig and a few weeks of honest cleanup. It is the best deal in software. ## What strict actually turns on The strict flag is a bundle. The one that matters most is strictNullChecks, which makes null and undefined separate types, so that a function returning a possibly missing user cannot be used as if the user were always there. Nearly every production null pointer I have traced back in an audit came from a codebase where this was off. noImplicitAny forbids the compiler from silently giving up and typing something as any. strictFunctionTypes makes callback parameters checked the right way around. useUnknownInCatchVariables types the caught error as unknown instead of any, so you have to check what it is before you read its message. strictPropertyInitialization insists that class fields are set in the constructor. Then there are the flags that are not in the bundle and should be. noUncheckedIndexedAccess makes an array or record lookup return a possibly undefined value, which is what it is. exactOptionalPropertyTypes distinguishes a property that is absent from one set to undefined. noFallthroughCasesInSwitch and noImplicitOverride catch two of the most common copy-paste mistakes. I turn all of them on, on day one, before there is code to complain. ## What it catches that reviewers miss The missing case. A switch over a union of order statuses that handles four of five. With strict checks and a `never` assertion in the default branch, adding a fifth status breaks the build everywhere the switch forgot it. A human reviewer sees a diff that adds a status and approves it. The compiler sees eleven switches that need updating. The optional that was not. An API response type marked a field as optional because one endpoint omits it, and now a component that uses it for display renders undefined. Without strict null checks that is a runtime blank. With them it is a red squiggle before the commit. The index that missed. Reading the first element of an array that came back empty. Reading a config map by a key that was never set. noUncheckedIndexedAccess makes every one of those a compile error unless you handle the miss, and it is uncomfortable for exactly one week, after which the code is honest about its assumptions. ## Turning it on in an old codebase The objection is always the count. Turning on strict in a mature project produces two thousand errors and nobody has a week. The answer is to not do it all at once. Enable one flag at a time, starting with noImplicitAny, then strictNullChecks. For each, suppress the existing errors with a directive that includes a ticket reference, and add a CI check that the count only goes down. New code is strict from the first day. Old code gets cleaned when it is touched. In most projects the count reaches zero in a quarter without a dedicated effort, because engineers fix the suppression when they are already in the file. Project references let you go further: strict packages and lenient packages in the same repository, with the boundary between them explicit, so that the domain core is strict even while the legacy admin panel is not. ## The review that remains Strict mode does not replace human review. It removes the mechanical part of it, and that changes what humans review. When the reviewer knows the compiler already checked nullability and exhaustiveness, they can spend their attention on the questions only a person can answer: is this the right design, does this name mean what it says, will we understand this in a year. That is the review worth paying a senior engineer for. The other kind costs one line in a config file, and I have never understood why anyone leaves it off. --- ## Capacity planning on a napkin 2025-10-28 · system-design Before I read a single line of architecture for a new system, I want to see the napkin. Not a spreadsheet, not a load test. A handful of numbers, multiplied together, that say roughly how big this thing is. If the team cannot produce it, the architecture is a guess dressed up as a plan. ## From daily numbers to per-second numbers Business people think in days and months. Systems live in seconds. The conversion is the whole trick. A day has 86,400 seconds. Round it to 100,000 and you can do it in your head: one million requests per day is about 10 per second on average. Ten million is 100 per second. That average is misleading, though, because traffic is not flat. A typical consumer product does 3x its average at peak; a B2B tool used during office hours in one timezone, more like 5x; something with a scheduled event, a sale, a TV mention, 10x or more. Design for peak, and write down which multiplier you chose. So: a clinic scheduling app with 2,000 clinics, each doing 300 bookings a day, is 600,000 bookings a day, about 7 per second average, maybe 35 per second at peak, and each booking might trigger 10 reads. That is 350 reads per second at peak. A single Postgres instance on decent hardware handles several thousand simple queries per second. This system does not need sharding. It needs indexes. ## Storage grows in a straight line, until it doesn't Storage is bytes per record times records, then times growth. A booking row with a few ids, timestamps, a status, and a note is perhaps 500 bytes with indexes. 600,000 a day is 300 MB a day, about 9 GB a month, 110 GB a year. Comfortable. Now add an audit log of every state change at 1 KB each, five changes per booking. That is 3 GB a day, 1 TB a year, and suddenly retention policy is an architectural decision, not an afterthought. That is the pattern: the main table is rarely the problem. Logs, events, attachments and anything with a fan-out are what fill disks. Do the multiplication for every table with a fan-out. ## Bandwidth and memory, same method Bandwidth: response size times requests per second. A 50 KB JSON payload at 350 rps is 17.5 MB/s, around 140 Mbps. Fine for a server, painful if it crosses a cloud egress bill. Memory: the working set, meaning the rows that are actually hot, should fit in RAM for the database to be fast. Today's bookings across all clinics are a tiny fraction of the 110 GB. That is why the system stays fast even as the table grows. ## The goal is the order of magnitude None of these numbers will be right. That is not the point. The napkin tells you whether you are building for 10 rps or 10,000, for 10 GB or 10 TB, and those are different systems with different budgets. Being wrong by 2x changes nothing. Being wrong by 100x means you picked the wrong architecture. So the most valuable thing on the napkin is not a number, it is knowing which number you are least sure about. Number of clinics? Bookings per clinic? The fan-out of reads per booking? Put a star next to it. That is the assumption to validate first, with real data, before you commit to infrastructure. ## Write it down, then come back Put the napkin in the repo. A markdown file with the assumptions, the multipliers and the date. Every six months, or after any surprise, compare it to the real metrics. The gap between what you assumed and what happened is the most honest architecture review you will ever get, and it costs an hour. --- ## Small models, big leverage 2025-10-25 · ai, system-design When a team shows me an AI pipeline, the first thing I count is how many calls go to a frontier model. Usually the answer is all of them. Classification, extraction, routing, reformatting, summarizing a paragraph, all sent to the biggest and slowest model available, because that is what worked in the prototype. The prototype had one user. Production has a bill, a latency budget and a queue. The fastest way to fix all three is to notice that most of those calls do not need a frontier model at all. ## Where small models win A small model is good at narrow, well-specified tasks with clear outputs. Is this ticket about billing or about login. Extract the order number from this email. Does this paragraph mention a side effect. Rewrite this in the house style. For tasks like these, a small model with a good prompt or a light fine-tune reaches the quality bar, runs in a fraction of the time, and costs a fraction per call. In pipelines I have restructured, the majority of calls by volume fell into this category. The frontier model was doing the equivalent of using a data center to add two numbers. ## Where they lose Small models fail on tasks that need broad world knowledge, long multi-step reasoning, or robustness to messy instructions. They follow a precise prompt well and a vague one badly. They are more sensitive to input length and formatting. And they are worse at knowing when they do not know, which matters when a wrong confident answer is expensive. So the design is not "replace the big model". It is "route by task": the small model handles the narrow, high-volume steps, and the frontier model handles the open-ended or high-stakes ones, with an explicit criterion for which is which. ## Cascades The pattern I use most is a cascade. The small model tries first. If its output passes validation and its confidence, however you measure it, is above a threshold, done. If not, the input goes to the frontier model. For most distributions of real inputs, the small model resolves the majority and the frontier model sees only the hard tail. Cost and latency drop roughly in proportion to what the small model catches, and quality on the hard cases is unchanged because those still reach the big model. The threshold is a tunable, and you tune it with your eval set, not with intuition. ## The eval set makes it safe None of this is responsible without a way to measure quality per task. Before I move a step to a small model I run the eval set for that step against both, and I look at where they disagree. Sometimes the small model is fine. Sometimes it is fine on ninety percent and catastrophic on a specific category, and the cascade needs a rule that sends that category straight to the frontier. The eval set turns a scary migration into an afternoon of table reading. ## Ownership Small models also change who controls the system. A model you can run yourself, on hardware you choose, is a dependency you can pin, version, and keep when the provider deprecates it. For pipelines that process sensitive data, running the narrow steps locally means that data never leaves your infrastructure for those steps, which shrinks the compliance surface. That alone has justified the engineering in more than one system I have designed. The big model is a great tool. Using it for everything is not a decision, it is the absence of one. --- ## A TypeScript monorepo that stays fast after year two 2025-10-23 · fullstack, typescript Every monorepo I audit in its second year has the same complaint: the type check takes four minutes, the editor lags, and a change in one package rebuilds everything. The team blames the size. The size is never the cause. The cause is that the dependency graph was never designed. Packages import each other freely, one shared package imports everything, and the compiler has no choice but to treat the whole repository as a single program. ## Project references are not optional TypeScript can compile a monorepo incrementally, but only if you tell it the shape. Project references, declared in each package's tsconfig, let `tsc --build` compile packages in dependency order and skip the ones whose inputs did not change. Without them, every type check starts from zero. The prerequisite is that every package has a real boundary: its own tsconfig, its own declared dependencies, and no imports that reach into another package's source through a relative path. The moment a file does `../../other-package/src/thing`, the boundary is gone and so is the incrementality. I enforce this with a lint rule that forbids relative imports across package roots, and with `isolatedModules` and `verbatimModuleSyntax` turned on, so that every file can be transpiled alone and type-only imports are explicit. Those two flags also keep the build compatible with the fast transpilers that do not type check at all. ## The shared package that eats the graph The most damaging pattern is the package named utils or common or shared. It starts with three helpers. In year two it exports database clients, React components, validation schemas and the pricing engine, and every other package depends on it. A change to one helper invalidates the entire repository. The fix is to split by rate of change and by audience. Types that everyone needs go in a package that changes rarely and has zero runtime dependencies. UI components go in a package that only web apps import. Domain logic goes in packages named for the domain. The rule I apply: if two things in a package have different reasons to change, they are two packages. Go further and check the direction of dependencies. Domain packages must not import from application packages. Types packages must not import from anything. A tool that draws the graph and fails the build on a cycle costs an afternoon to set up and saves months. ## Caching what the compiler already knows Incremental builds keep a .tsbuildinfo file with the previous state. In CI, that file must be restored from cache or the build starts cold on every run. I have seen teams with perfect project references and a fifteen-minute CI type check because the cache key included the commit hash and never hit. A task runner with remote caching takes this across machines. If a package's inputs match a previous build anywhere in the team, the outputs are downloaded instead of rebuilt. That is the single change that makes year two feel like year one, and it depends entirely on the boundaries being real, because the cache key is the package's inputs. Type checking and bundling are separate concerns. The bundler should transpile without type checking, and the type check should run as its own step, per package, in parallel. Coupling them means the slowest possible path on every build. ## The editor problem Editor lag is a different bottleneck. The language server loads one program per tsconfig it can find, and if the root tsconfig includes every package, it loads the whole world. Keep the root config minimal and let each package's config be the unit the editor opens. Solution-style configs with references and no files of their own exist for exactly this. The native TypeScript compiler port changes the constants here, and it is a real improvement, but it does not change the shape of the problem. A graph with no boundaries is still one program, and one program is always the slowest thing you can ask a compiler to check. Draw the graph on day one. It is the cheapest hour of the whole project. --- ## Conway's law is a tool, not a curse 2025-10-21 · architecture, career Conway's law says that a system's structure will mirror the communication structure of the organization that builds it. Most engineers quote it as a lament: 'of course the API is a mess, look at how the teams are split'. I read it the other way. If the structure of the system follows the structure of the teams, then team structure is an architecture tool, and I can use it. This is not a management topic. It is the most powerful refactoring instrument available to an architect, and it does not require touching a line of code. ## The law works in both directions The usual reading is that teams shape systems. Three teams build three services, and the service boundaries land exactly where the teams stop talking. That part is well documented and I see it in every audit: the seams in the code match the seams in the meeting invites. The less discussed direction is that systems shape teams. Once the three services exist, they define who talks to whom. A new hire joins 'the payments service team', and the boundary is now self-reinforcing. Even if the original split was wrong, the organization has adapted to it, and every person's job title depends on it staying. This is why architecture changes fail when they are attempted as pure code changes. You cannot merge two services into one module while the two teams still have separate backlogs, separate on-call rotations and separate managers. The code will drift apart again within two quarters, because the organization that produces it has not changed. ## The inverse maneuver The practical move is called the inverse Conway maneuver: decide what system structure you want, then arrange the teams to produce it. If you want checkout and payments to be one cohesive module with one contract to the outside world, put the people who build them in one team with one backlog. The code will follow, not because of a mandate, but because the daily conversations now cross the old boundary. I have recommended this more often than I have recommended a rewrite, and it works more often. A rewrite fights the organization. A reorganization enlists it. It also exposes fake boundaries. If two services can only be maintained by the same three people, they are one service that happens to deploy twice. If a 'platform team' owns a shared library that every product team modifies weekly, the library is not a platform, it is a coordination bottleneck, and the org chart is lying about ownership. ## Where I see it go wrong The most common failure is splitting teams by technical layer. A frontend team, a backend team, a data team. Conway's law then produces a system with three layers that must all change for every feature, and three handoffs per feature. The result is a system whose boundaries protect nothing, because no business change lives inside one of them. The second failure is splitting by headcount rather than by domain. A team grows to twelve, so it splits into two teams of six, and the split is drawn wherever it is politically easiest. The resulting service boundary has no domain meaning. Six months later, the two teams are in every meeting together because every change crosses the line. The third is ignoring the law when hiring contractors or agencies. An external team that builds a module in isolation produces a module that is isolated: its own conventions, its own error handling, its own idea of what a customer is. That is Conway's law again, and it is predictable. If you want the module to integrate, someone from the core team has to be in the room every week. ## The questions I bring to leadership When I audit a system and find a structural problem, I ask about the teams before I propose a code change. Who owns this boundary? Who has to agree for it to move? Which two teams argue most often, and is the argument about a line that should not exist? Then I draw two pictures: the system as the code shows it, and the system as the org chart implies it. Where they disagree, that is where the pain is, and it is usually where the next incident will be. The fix is often uncomfortable, because it means changing who reports to whom, which is a harder conversation than which folder a file goes in. But it is the fix that lasts. Draw the org chart on purpose, and the architecture will follow it. It was always going to. --- ## The pull request that looks fine 2025-10-19 · audit, engineering The dangerous pull request is not the thousand-line refactor. Everyone reads that one carefully, or refuses to merge it at all. The dangerous one is forty lines, has a clear title, passes every test and gets approved in eight minutes by someone who trusts the author. I have traced dozens of incidents back to a change like that, and they share an anatomy. ## It changed a default The most common shape: a function that had a parameter with a default, and the default changed. Timeout from 30 seconds to 5. Page size from 50 to 500. A boolean flag that used to be false now true. The diff is one line. Every caller that never set the parameter explicitly, which is most of them, just changed behavior without appearing in the diff at all. The tests passed because the tests set the parameter. Production did not. ## It touched something shared A helper in a utils file, a base class, a middleware, a database model that six features read from. The change was correct for the feature the author was working on. It was wrong for two of the other five, and neither of them had a test that would notice, because their tests mock the shared thing. The reviewer read the diff and saw a reasonable change to a helper. Nobody read the callers, because the tooling shows you the diff, not the blast radius. ## It removed a check that looked redundant A null check that 'can never happen'. A validation that 'the frontend already does'. A guard around a retry that seemed paranoid. Someone cleaned it up in passing, in a pull request about something else, and the reviewer nodded because the code was cleaner. The check was there because of an incident two years ago that nobody on the current team remembers. The incident came back. ## It changed something the diff cannot show The change added a column and started writing to it in the same deploy. Fine on the developer's machine, where the migration and the code start together. In production the code rolled out to three instances before the migration finished, and for ninety seconds every write failed. Or the reverse: the migration dropped a column the old code still read, and the old code was still running during the rollout. The diff was correct. The order was not, and order is invisible in a diff. Serialization is the other invisible one. A field renamed in an API response. A date that used to be a string is now an ISO timestamp. An enum value spelled differently. A number that was an integer is now a decimal. The backend tests passed because they test the backend. The mobile app that was released three months ago, and cannot be updated on every phone tonight, parsed the old shape. This one is the quietest, because it often fails only for some clients, and the errors show up on their side, not yours. ## How I review the one that looks fine I have stopped trusting my sense that a change is small. Instead I ask, mechanically, for every pull request regardless of size: did a default change? Is anything touched here imported from more than one place? Was anything removed, and does the description say why it was safe? Does the deploy order matter? Does any shape that leaves the system change? It takes five minutes and it is boring. The pull requests that fail those questions are almost always the ones that would have looked fine. Those five minutes are the whole difference between a review and an approval. An approval says the diff is reasonable. A review says the system will still work after it. --- ## Backups you have never restored are hypotheses 2025-10-17 · infra, reliability I ask the same question in every audit: when did you last restore a backup, to a real environment, and use the result? The most common answer is a pause. The second most common is 'the platform does it automatically'. Both answers mean the same thing. The team has a hypothesis that they can recover, and they have never tested it. A backup that has never been restored is not a backup. It is a file that might be a backup. The difference is only discovered on the day it matters, which is the worst possible day to discover anything. ## The ways restores fail The backup is there but the credentials to decrypt it were rotated and the old key is gone. The backup is there but it is of the database only, and the object storage with the uploaded files was never included. The backup is there but it is four hundred gigabytes and restoring it takes eleven hours, while the business assumed one. The backup is there but the schema it contains is three migrations behind the code that is now deployed, and nobody knows which version of the code matches it. The backup is there but the restore procedure references a server that was decommissioned. The backup is there but it was taken mid-transaction with an inconsistent state that the database refuses to load. The backup is there but it has been silently empty for six weeks because a cron job failed and no alert was set on its absence. Every one of these is something I have seen, and none of them would have survived a single rehearsed restore. ## What a real restore test looks like Take the most recent backup. Restore it to a fresh environment that has nothing in common with production except the software versions. Start the application against it. Log in. Run the three most important user flows. Compare a few known records against production. Time the whole thing with a stopwatch. Write down every step someone had to improvise. That last item is the payoff. The improvised steps are the ones that will be forgotten during the real incident. Each one goes into the runbook, and the next rehearsal should have fewer of them. When the restore can be executed by someone who did not write the runbook, from a clean laptop, within the time the business expects, then you have a backup. ## Two numbers you must decide Recovery point objective is how much data you are willing to lose, measured in time. If backups are nightly, the answer is up to twenty-four hours, and you should say that sentence out loud to whoever owns the business. If that is unacceptable, you need continuous archiving of the write-ahead log, or replication, and those are different investments with different failure modes. Recovery time objective is how long the business can be down while you restore. This number is decided by the business and constrained by physics. A large database cannot be restored from a dump in minutes. If the objective is minutes, the answer is a standby replica, not a faster restore. Neither number can be chosen by engineering alone. Both must be written down, and the backup strategy must be checked against them, not against what the platform offers by default. ## The minimum discipline - Backups cover every stateful thing: database, object storage, secrets, configuration, and the queue if it holds unprocessed work. - An alert fires when a backup does not complete, not only when it fails loudly. - At least one copy lives in a different account or provider, with credentials that a compromised production account cannot reach. - A restore is rehearsed on a schedule, quarterly at minimum, and the measured time is compared to the objective. - The runbook is updated after every rehearsal, by the person who did it. It is a few hours per quarter. The alternative is learning, during an outage, that the thing you paid for every month was a hypothesis all along. --- ## Schema evolution without downtime 2025-10-15 · system-design, data The most dangerous line in a deployment is not in the application code. It is ALTER TABLE. Schema changes touch the one component that cannot be restarted casually, cannot be rolled back with a redeploy, and holds every customer's data at once. Yet most teams treat migrations as a formality: write the migration, run it in the pipeline, ship. That works right up until the table has 80 million rows and the migration takes a lock at 2 pm on a Tuesday. ## Expand, then contract The core discipline is expand and contract. Never make a change that the running application cannot survive. Instead, split every change into a phase that only adds and a phase that only removes, with a period in between where both old and new shapes coexist. Take the classic case: renaming a column from "name" to "full_name". The one-step rename breaks every running instance the moment it executes. The expand and contract version has five steps. Add full_name as nullable. Backfill it in batches of a few thousand rows, sleeping between batches. Switch the application to dual-write, meaning every write updates both columns. Switch reads to full_name once the backfill is complete and verified. Only then, after a deploy or two of confidence, drop name. It is more work. It is also the only version that never has a moment where the database and the code disagree about the world. ## Overlapping versions are the normal case The reason expand and contract exists is that deploys are not atomic. During a rolling deploy, version N and version N+1 run side by side for minutes, sometimes longer. If the schema only works with one of them, you have downtime during every deploy and, worse, a rollback of the code becomes impossible because the old code no longer understands the schema. Every migration should be tested against the previous application version too. If the old code cannot run against the new schema, you have not finished the expand phase. ## Locks and long-running operations Some operations that look innocent are not. Adding a NOT NULL column with a default used to rewrite the entire table on older Postgres versions, holding an exclusive lock the whole time. Modern versions handle this cheaply, but adding a NOT NULL constraint to an existing column still needs a full scan. Adding a foreign key validates every row. Creating an index locks writes unless you create it concurrently, which cannot run inside a transaction. The rule I follow: any migration on a large table gets read as a question. What lock does this take, and for how long? If the answer is "I am not sure", it does not ship until someone is sure, ideally by running it against a production-sized copy and timing it. ## Events have schemas too Everything above applies to message contracts. An event on a queue is a schema that consumers depend on, and consumers do not redeploy in lockstep with producers. Add fields, never remove or rename them in one step. Tolerate unknown fields on the consumer side. Version the event when the shape genuinely changes, and keep consuming the old version until every producer has moved. ## Flags and rollback plans For the cutover steps, switching reads and switching writes, use a feature flag rather than a deploy. A flag flips back in seconds. A deploy takes minutes and a nervous engineer. And before any contract phase, write the rollback plan as literally as a runbook: if we drop this column and something breaks, what do we do? If the honest answer is "restore from backup", wait longer before dropping. Keeping an unused column for another month costs nothing. Losing data costs everything. Schema evolution without downtime is not clever. It is slow, boring and split into small steps. That is exactly why it works. --- ## Structured outputs are contracts 2025-10-11 · ai, engineering The single change that most improved the reliability of the LLM features I ship was not a better prompt or a better model. It was refusing to accept free text. Every call that feeds a program returns a schema: typed fields, enumerated values, explicit nulls. The moment the output has a shape, the model stops being a chat partner and becomes a component with a contract, and everything I know about contracts applies. ## A schema is a promise When I define an output schema, I am making a promise to the rest of the system: these fields will exist, these types will hold, these enums will be one of these values. That promise is the integration point. The code downstream is written against the schema, not against the model, which means I can change the prompt, swap the provider or upgrade the model, and as long as the schema is honored, nothing downstream notices. That is the same benefit an API contract gives between services. The model is just a service with a very unusual implementation. ## Validate, do not trust Asking for a schema is not the same as getting one. Even with a provider that enforces the shape, the values inside it are still generated: a confidence field will be filled with a plausible number, an enum will be chosen even when none fits, a date will be well-formed and wrong. So every structured output passes through the same validation layer as any external input: parse, check types, check ranges, check referential integrity against the source. When validation fails, the failure is handled explicitly, with a retry, a fallback or an error, never with a silent default. I treat the model exactly as I would treat a third-party API that is usually right. ## Design the schema for the consumer, not the model A common mistake is to design the schema around what the model finds easy to produce. Design it around what the code needs to consume. If the downstream logic needs to know whether a fact was found in the source, give it a boolean and a citation field, not a free-text explanation. If a decision has three outcomes, give it an enum with three values and a fourth for "cannot determine", because the absence of an escape hatch is how a model is forced to guess. Make required fields required and optional fields optional. The schema is where you encode the invariants that the prompt can only suggest. ## Versioning and evolution Schemas change, and a schema used by a model changes for two reasons: the product needs a new field, or the model does better with a different shape. Both are migrations. I version output schemas the way I version API responses: additive changes are safe, removals and renames get a transition period, and the eval set is re-run against the new shape before it ships. The prompt and the schema travel together in version control, because a prompt written for one schema will quietly degrade against another. ## What breaks in practice The failures I find in audits are consistent. Deeply nested schemas the model fills inconsistently, when a flat one would have worked. Fields that mean different things in different prompts because nobody wrote the definition down. Validation that exists in one code path and not in the batch job that uses the same model. And the classic: a schema that was never enforced, so the code parses free text with a regex and hopes. Every one of those is a contract problem with a known solution. The model made it visible. The engineering fixes it. --- ## Timeouts are the cheapest resilience you will ever buy 2025-10-09 · system-design, reliability Of everything I check when I audit a system, timeouts are the item most likely to be missing and the cheapest to fix. One line of configuration, sometimes one argument in a function call, and the difference between "one dependency is slow" and "the whole platform is down." The default in most HTTP clients, database drivers and message consumers is no timeout at all. No timeout does not mean "reasonable." It means infinite. Your code will wait forever for an answer that is never coming, holding a connection, a thread, and a slot in a pool while it does. ## What actually breaks A slow dependency does not fail loudly. It just answers later and later. Each request that waits holds resources. Your worker pool has maybe 50 threads, or your connection pool has 20 slots. Once every one of them is waiting on the slow thing, requests that have nothing to do with it start queueing behind them. Health checks fail. The load balancer marks you unhealthy. Now you are down because a recommendation widget was slow. That is thread pool exhaustion, and the timeout is what prevents it. A hung call that gives up after two seconds returns its thread to the pool. A hung call with no timeout keeps it until the process restarts. ## Three numbers, not one A timeout is not one number. At minimum you need to think about three. The connect timeout covers establishing the connection. If the host is unreachable or overloaded, you learn it here. This should be short, often a few hundred milliseconds inside a data center, a second or two across the internet. The read timeout covers waiting for bytes once connected. This is where most hangs live. It depends on what you are calling, but if your own users are waiting on this, it should be well under what they would tolerate. The total budget covers the entire operation including retries. This is the number that matters to the user. A one second read timeout with three retries is a four second operation. Design that on purpose, not by accident. Illustrative values, not rules: a cache lookup might get 50 milliseconds, an internal service call 500 milliseconds to a second, a third-party payment API a few seconds because you have no choice, and a batch job minutes. The point is that each one is a decision. ## Deadlines travel with the request A user request arrives with, say, three seconds of patience. Your API calls service A, which calls service B, which queries a database. If each of them independently decides on a three second timeout, the user can wait nine seconds for a failure. The right pattern is deadline propagation: the entry point picks a deadline, passes the remaining time down each hop, and every callee uses a timeout shorter than what it was given. The rule that follows: your timeout must be shorter than the timeout of whoever is calling you. If the caller gives up at two seconds and you keep working for five, you are burning resources on an answer nobody will read. ## Pair with retries, carefully A timeout without a retry is a fast failure. A timeout with an unbounded retry is a self-inflicted denial of service. Pair them deliberately: a small retry count, exponential backoff, and jitter so a thousand clients do not retry in the same millisecond. Only retry operations that are idempotent, or that you have made idempotent with a key. And when the timeout fires, do something useful. Return cached data, degrade the feature, show a partial page. The timeout is what gives you the chance to make that choice. Without it, the choice is made for you by the slowest dependency you have. I have never seen a mature system with too many timeouts. I have seen plenty go down for lack of one. --- ## Open science needs boring infrastructure 2025-10-05 · science, infra Everyone in science agrees with open science in principle. Share the data, share the code, let others check and build. Then you look at where the shared data actually lives: a lab website on a university server that will be reorganized next year, a cloud bucket on a grant that ends in eighteen months, a link in a paper that already returns 404. The ideal is fine. The infrastructure is what is missing, and it is missing because it is boring. ## The boring list Persistent identifiers that resolve in ten years. Storage with a funding model that outlives the grant. Backups that someone has actually restored. Formats that a program can read without a vendor license. An API, so that people can fetch data without clicking through a portal. Access control for the data that cannot be fully open. Monitoring, so somebody knows when the thing goes down before a reviewer does. A person, paid, whose job it is. None of that gets a paper written. All of it is required for the paper to remain checkable. ## What I have seen break A shared dataset that lived as a zip on a personal drive. When the postdoc left, the drive went with them. A code repository that pinned nothing, so three years later nothing installed. A public database where the schema changed without a version and every downstream script silently produced different numbers. An identifier scheme that was a spreadsheet row number, meaning every re-sort reassigned identities. I have watched businesses lose money to each of these patterns and fix them within a quarter, because money makes infrastructure urgent. Science has no equivalent forcing function, so the same failures persist for a decade. ## Boring is a feature In production engineering, the highest compliment for infrastructure is that nobody thinks about it. Postgres is boring. Object storage is boring. Plain CSV with a documented schema is boring. Every one of those will still work in fifteen years, and nobody will need to be an expert to use them. The temptation in research software is to build something clever, a custom platform with a custom format and a beautiful portal. The clever thing needs a maintainer, and maintainers leave. The rule I apply is the same one I use for clients: choose the most boring tool that meets the requirement, and spend your creativity on the science. ## The economics nobody funds Infrastructure is a cost center with no publication attached. Grants pay for discovery, not for the server that keeps last year's discovery accessible. This is a structural problem and I do not have a policy fix for it. What I do have is an engineering fix: make the infrastructure so cheap and so standard that its cost rounds to zero. A well-organized bucket with checksums and a README is nearly free. A bespoke portal is not. The way to make open science sustainable is to lower its operating cost until nobody has to argue for it. ## What an engineer can offer I spend my days making systems that stay up and stay honest for businesses. The same skills apply here without modification. Versioned data, immutable raw layers, reproducible environments, uptime monitoring, documented interfaces. It is not glamorous work and it is the work I am best at. If open science is going to be more than a slogan, someone has to lay the pipes, and I would rather that person understand what happens when pipes fail. Every advance needs validation, and validation needs the data to still be there when someone goes looking. Boring infrastructure is how it stays there. --- ## Technical debt is a loan, and every loan has a name on it 2025-10-02 · architecture, career The technical debt metaphor is older than most of the codebases carrying it, and it has been worn smooth. Teams say 'we have a lot of tech debt' the way they say 'the weather is bad': a condition, nobody's fault, nothing to do but complain. That is not how loans work. A real loan has a borrower, a principal, an interest rate and a due date. When any of those is missing, it is not a loan. It is a gift, a theft or a mess. Most technical debt I audit is a mess, and the fix is to give it back the things a loan has. ## The name Every shortcut was taken by someone, for a reason, under a pressure. The commit exists. The pull request had a reviewer. When a team says 'nobody knows why this is like this', what they mean is that the borrower left, or was never written down. So the first rule: debt gets a name at the moment it is taken on. Not to blame, but to keep the context. A comment, a ticket, a line in an ADR: 'skipped retries on the payment client to hit the launch date, owner is this team, revisit after the first month of traffic'. The person who did it is the one who knows what the safe version looks like. Debt with a name gets paid. Debt without a name gets inherited, and heirs do not pay debts whose terms they cannot see. ## The interest Not all debt costs the same. A missing test on an internal admin page costs almost nothing per month. A hand-rolled auth check duplicated in twenty route handlers costs every time someone adds a route and forgets it, and it pays out as a security incident. Interest is the recurring cost: the extra hour on every feature that touches the area, the incident every quarter, the onboarding week lost to explaining a workaround. I ask teams to estimate it in hours per month, roughly. The number is always wrong and always useful, because it turns 'we should refactor this' into 'this costs us a day a week', which a product owner can weigh. The high-interest debt is usually not the ugliest code. It is the shortcut that sits on the main path: the data model that forces every query to join five tables, the missing idempotency on the endpoint every client retries, the deploy that requires a manual step. Pay those first, regardless of how much the ugly code bothers you. ## The term Some debt is meant to be permanent. A prototype's shortcuts are fine if the prototype is thrown away. A migration shim is fine if the migration finishes. The failure is when temporary things have no end date, and the shim becomes load-bearing. So every piece of named debt gets a term: a date, a milestone or a trigger. 'Until we pass ten thousand daily orders'. 'Until the second tenant signs'. 'Until next quarter's planning'. When the term arrives, the debt is either paid or consciously refinanced with a new term, by someone who can say why. ## The ledger With names, interest and terms, debt can be listed. I keep a single ledger per system: one page, each entry with what was skipped, why, who owns it, the estimated monthly cost and the term. Twenty lines is typical. Two hundred means the ledger is being used as a wish list, and wish lists are not debts. The ledger changes conversations. Instead of 'we need a refactoring sprint', which no product owner has ever loved, the ask becomes 'these three entries cost us about six days a month and have passed their term, here is the two-week plan to close them'. That is a business case, and it usually wins. ## What this asks of engineers It asks for the discomfort of putting your name next to a shortcut. Engineers avoid that because it feels like admitting fault. It is the opposite. Signing a shortcut says: I understood the trade-off, I made it on purpose, and I know what paying it back looks like. That is the mark of someone who can be trusted with the next one. The debt is not the problem. Unnamed debt, with no interest estimate and no term, is the problem, because it cannot be reasoned about and therefore never gets paid. Give every loan a name and the backlog starts to move. --- ## From the print shop to the whiteboard 2025-10-01 · career, life There is a plotter in a print shop in São Paulo that taught me more about architecture than any book. I was about twelve. Samuel, who ran the shop, put me on CorelDRAW and vinyl cutting, and the first thing I learned was that the machine does exactly what the file says. Not what you meant. What you said. I want to trace the line from that room to the whiteboard where I design systems today, because the line is straighter than it looks, and because I think many engineers underrate what they learned before they wrote code. ## Vinyl does not forgive A vinyl cut is a one-way operation. You send the vectors, the blade follows them, and if a path was not closed or a node was off by a millimetre, the letter falls apart when you weed it. There is no undo. There is a roll of wasted material and a client waiting. So you learn to check the file. You zoom in on every join. You check that the text was converted to curves so the plotter does not substitute a font. You do a small test cut before the big one. Twenty years later this is exactly how I treat a database migration. Migrations are one-way. Production does not forgive. You run it on a copy first, you check the joins, you make sure nothing gets substituted silently, and you do the small one before the big one. ## Electronics taught me to read a system After the print shop, a man named Val taught me electronics repair. A board arrives dead. You do not know why. There is no error message. What you have is a schematic, a multimeter, and a way of thinking: power goes in here, it should come out there, so where does it stop? That is how I audit software now. I follow the request from where it enters to where it should leave, and I find where it stops. The tools are different, the discipline is the same. You do not guess. You measure at each stage. And you learn to distrust the component that "cannot be the problem", because on a repair bench that component is the problem about half the time. ## Painting and construction taught me sequence As a teenager I paid my own rent working in auto parts, events, painting and construction. Construction has an order and the order is not negotiable. You cannot paint before the plaster dries. You cannot wire before the walls are up. If you rush the sequence to hit a date, the finish cracks and you do it twice. Software has the same sequence, and teams violate it constantly. They build the UI before the data model is settled. They add the cache before the query is correct. They ship the feature before the failure mode is designed. On the whiteboard, the first thing I draw is the order in which things must be true, and only then the boxes. ## What the whiteboard actually is People think the whiteboard is where the clever part happens. It is not. The whiteboard is where I do the test cut. It is where I trace the current through the system before anyone builds it. It is where I check the sequence so that the finish does not crack. The habits are the same ones I had at twelve. Check the file before you send it. Follow the signal until it stops. Respect the order. The only thing that changed is that the material got more expensive. A wasted roll of vinyl cost Samuel an afternoon. A wrong architecture decision costs a company a year. ## For the engineer who came from somewhere else If you came to software from a trade, a shop floor, a kitchen, a construction site, do not treat that as time lost before your real career started. It was your real career. You learned to respect irreversible operations, to trace a fault instead of guessing, and to do things in order. Most computer science graduates spend their first five years learning those three things the expensive way. Put them on your whiteboard. They are worth more than you think.