# phavella · Custom AI, built for production. · full reference

> phavella is an AI solutions studio. We design, build, and integrate custom AI into the stack a company already runs: agents, apps, voice, document intelligence, analytics, decisioning, and creative pipelines. Use cases big and small, from a single workflow to a full platform. A human signs off on every consequential decision; the routine majority auto-executes. Then we stay until it works in production.

- **What we are:** A senior, hands-on AI solutions studio: four named principals who build and run the systems themselves. No juniors on production systems.
- **What we build:** Custom AI, end to end: full-stack products, multi-agent systems, voice interfaces, retrieval at scale, dashboards, decision systems, and creative pipelines. We build from scratch or integrate into the stack you already run (CRM, inbox, docs, ads, ops).
- **How we work:** Scope, Build, Ship, Operate. Weekly demos on your data, not slides. Evals and monitoring wired in. Routine work auto-executes; decisions that are expensive to get wrong route to a named human reviewer with context and an audit trail.
- **What we won't do:** Let AI make a consequential decision alone. Rip out your system of record. Free pilots. Equity-only deals. Strategy decks without a build behind them. Engagements where the math doesn't work.

## The nine practices · in full

### 1. AI Agents & Workflow Automation

- **URL:** `https://phavella.com/capabilities/agents-automation`
- **Tagline:** Multi-agent systems and automations that take routine work off your team's plate.
- **Summary:** Agent systems that execute real workflows (research, routing, reporting, operations) with human sign-off exactly where it counts. 618+ automations shipped across 12 industries.
- **Delivers:** multi-agent orchestration; n8n/workflow automation; MCP servers & tool integrations; approval-gated execution; agent monitoring & evaluation
- **Proof:**
  - 618+ · automations shipped (PARIK · 12 INDUSTRIES)
  - 34 · tools in one MCP server (METAADS-MCP)

### 2. AI-Powered Apps & Products

- **URL:** `https://phavella.com/capabilities/ai-apps`
- **Tagline:** Full-stack products with AI at the core, designed, built, and run end to end.
- **Summary:** From scope to production: applications where AI is the product, not a bolt-on. Our ad-production platform runs a 13-state production pipeline; our principal built 15 production AI apps at Newsweek.
- **Delivers:** product scoping & UX; full-stack build (Next.js / FastAPI / Postgres); AI feature engineering; deployment & operations; iteration against real usage
- **Proof:**
  - 15 · production AI apps (NEWSWEEK · 5 YEARS)
  - 13 · pipeline states in one product (IN-HOUSE AD PLATFORM)

### 3. Conversational & Voice AI

- **URL:** `https://phavella.com/capabilities/conversational-ai`
- **Tagline:** Assistants, copilots, and voice interfaces your customers and teams actually use.
- **Summary:** Chat and voice systems grounded in your data, with guardrails and human escalation: a financial chatbot across 4 agent frameworks at a wealth-management fintech, 20 medical AI agents at a clinical-intelligence company, production voice pipelines in our ad-production platform.
- **Delivers:** customer-facing assistants; internal copilots; voice synthesis pipelines (ElevenLabs); RAG-grounded answers; guardrails + human escalation
- **Proof:**
  - 20 · medical AI agents in production (CLINICAL-INTELLIGENCE CO)
  - 4 · agent frameworks in one chatbot (WEALTH-MGMT FINTECH)

### 4. Document Intelligence & Enterprise Search

- **URL:** `https://phavella.com/capabilities/document-intelligence`
- **Tagline:** Search, extraction, and understanding across millions of documents.
- **Summary:** Retrieval systems at real scale: a 30M-document Elasticsearch index at Newsweek, financial OCR for PE at a private-markets data platform, 66GB of retail catalogs at a construction-commerce marketplace, 100k+ clinical codes at a clinical-intelligence company.
- **Delivers:** RAG pipelines; vector + graph search; OCR & multimodal extraction (PDF, image, video frames); knowledge bases; retrieval-quality eval harnesses
- **Proof:**
  - 30M · documents in one search index (NEWSWEEK)
  - 66GB · retail catalogs processed (CONSTRUCTION-COMMERCE MKT)

### 5. Data, Analytics & Dashboards

- **URL:** `https://phavella.com/capabilities/data-analytics`
- **Tagline:** Pipelines, executive dashboards, and analytics that answer business questions.
- **Summary:** ETL, BI, and predictive analytics for operators: pharma analytics at a life-sciences commercial-analytics firm for three top-10 pharmaceutical companies; portfolio operations for two North American private-equity firms; embedded data work at a Fortune-tier technology company and a major airline.
- **Delivers:** ETL & data pipelines; executive dashboards; natural-language analytics; forecasting; data-quality monitoring
- **Proof:**
  - 3 · top-10 pharma clients served (LIFE-SCIENCES ANALYTICS)
  - 1+yr · embedded at a Fortune-tier tech company (FORTUNE-TIER TECH)

### 6. Risk, Fraud & Decision Systems

- **URL:** `https://phavella.com/capabilities/decision-systems`
- **Tagline:** Scoring, underwriting, and decisioning that survives regulators and scale.
- **Summary:** Production decision systems with model-validation and fair-lending discipline: a $150M risk portfolio and $1B annual transaction policy at Drip Capital; 16 million fraud decisions a month at Citibank.
- **Delivers:** fraud & risk models; underwriting / KYC / AML flows; model validation (KS, PSI); policy design; decision audit trails
- **Proof:**
  - 16M · fraud decisions per month (CITIBANK · 5+ YEARS)
  - $1B · annual transaction policy designed (DRIP CAPITAL)

### 7. Content & Creative Automation

- **URL:** `https://phavella.com/capabilities/creative-automation`
- **Tagline:** Generative pipelines for video, ads, and content, at production volume.
- **Summary:** Our ad-production platform produces video ads end to end (scripting, generation across 11 video models, voiceover, human approval gates) at 5× the weekly output. Creative OS shipped a 522-commit client ad platform.
- **Delivers:** video generation pipelines; ad creative automation; voiceover & audio; brand-safe approval workflows; multi-model routing
- **Proof:**
  - 5× · weekly ad output (IN-HOUSE AD PLATFORM)
  - 11 · video generation models orchestrated (IN-HOUSE AD PLATFORM)

### 8. System Integrations & AI Modernization

- **URL:** `https://phavella.com/capabilities/integrations-modernization`
- **Tagline:** AI wired into the stack you already run: CRM, inbox, docs, ads, ops.
- **Summary:** We integrate into what exists instead of replacing it: Slack, Google Drive, Sheets, Docs, Dropbox, and Meta Ads integrations shipped inside our ad-production platform; a 34-tool MCP server for Meta Ads; embedded modernization inside a Fortune-tier technology company and PE operating companies.
- **Delivers:** API & MCP integrations; AI layers over legacy stacks; workflow rewiring; data sync; incremental modernization roadmaps
- **Proof:**
  - 34 · tools in one Meta Ads MCP server (METAADS-MCP)
  - 7 · Slack channels orchestrated in one product (IN-HOUSE AD PLATFORM)

### 9. Evaluation & Reliability

- **URL:** `https://phavella.com/capabilities/evaluation-reliability`
- **Tagline:** The discipline that makes AI safe to run in production.
- **Summary:** Eval harnesses, regression suites, and monitoring so your AI behaves tomorrow the way it did in the demo: BFCL, TAU-bench, and DeepEval harnesses, multi-LLM comparison, hallucination reduction. Built into everything we ship; available standalone.
- **Delivers:** eval harnesses; multi-LLM benchmarking; hallucination reduction; monitoring & alerting; CI gates for AI behavior
- **Proof:**
  - 3 · public eval harnesses run (BFCL · TAU-bench · DeepEval) (AI REASONING-INFRA CO)

## The bench · in full

Four named, senior principals. The edge is the systems; the people are the proof.

### Parik Ahlawat · Principal · Agents & Orchestration

- **URL:** `https://phavella.com/bench/parik`
- **Headline:** Multi-agent platforms, model routing, n8n orchestration. 618+ automations shipped across 12 industries.
- **Skills:** multi-agent orchestration, agent orchestration, swarm, agents, mcp, model context protocol, tool calling, tool use, model routing, cost optimization, rag, agentic rag, langgraph, langchain, reasoning models, srm, small reasoning models, chain of thought, evaluation, bfcl, tau-bench, livebench, deepeval, state machines, fsm, fault tolerance, retry, fallback, video generation, video ad production, fal.ai, elevenlabs, meta ads, creative automation, full stack ai, fastapi, nextjs, typescript, react, kronus, production ai, mid-market, startup ai, telegram bot, slack, browser automation
- **Industries:** media, advertising, creative, video production, fintech, trading, crm, sales, developer tools, devtools, ai infrastructure, marketing saas, logistics, security, e-commerce, ecommerce
- **Tools:** n8n, apify, qdrant, pinecone, pgvector, chromadb, gemini, fal.ai, elevenlabs, meta ads api, shopify, zoom, slack sdk, telegram bot api, stripe, plaid, read.ai, zerodha, docker, kubernetes, gcp, aws, pm2, nginx
- **Engagements:** Kronus (own agent OS), Ad-production platform (in-house B2B build), An AI reasoning-infrastructure company (LLM evaluation), A smart-contract developer-tools company (code generation), Creative OS (522-commit client ad platform), metaAds-mcp (34-tool MCP server), An edge-AI company

### Arpit Jhamb · Principal · Data & Portfolio Ops

- **URL:** `https://phavella.com/bench/arpit`
- **Headline:** PE-backed data + portfolio consulting. ETL, executive dashboards, business automation across operating companies already running on real systems.
- **Skills:** data engineering, etl, etl pipelines, data analytics, executive dashboards, business intelligence, reporting, predictive analytics, business automation, ai agent workflows, agent engineering, lead qualification, lead scoring, segmentation, marketing automation, voice agents, image generation pipelines, pharma analytics, marketing mix modeling, mmm, multivariate regression, territory alignment, k-means, clustering, salesforce, crm integration, monitoring, observability, fairness, bias, responsible ai, governance
- **Industries:** private equity, pe, mid-market, portfolio, pharma, pharmaceutical, life sciences, b2c, e-commerce, ecommerce, marketing, public policy, government affairs, airlines, travel
- **Tools:** n8n, zapier, make, temporal, langchain, autogpt, pinecone, neo4j, faiss, chroma, mistral, llama, looker studio, powerbi, tableau, bigquery, gcp, aws, azure, docker, kubernetes, salesforce, hubspot
- **Engagements:** A North American private-equity firm, A mid-market private-equity firm, A Fortune-tier technology company (government affairs, 1+ yr embedded), A life-sciences commercial-analytics firm (three top-10 pharmaceutical companies), A major airline, Consumer-marketing and media brands

### Nikhil Bery · Principal · RAG & Document AI

- **URL:** `https://phavella.com/bench/nikhil`
- **Headline:** RAG, document intelligence, multi-LLM pipelines, vector + graph search at scale. 30M-doc Elasticsearch indexes.
- **Skills:** rag, retrieval augmented generation, vector search, document intelligence, document extraction, ocr, pdf extraction, pdf parsing, medical pdf, financial documents, embeddings, reciprocal rank fusion, hybrid search, bm25, knn, semantic search, multi-agent systems, multi-llm, multi-model, knowledge graph, entity extraction, spert, relation extraction, ner, named entity recognition, nlp, fine-tuning, gpt-4o fine tune, classification, clustering, umap, hdbscan, web scraping, crawling, anti-detection, selenium stealth, computer vision, detectron2, biomedical, biomedical knowledge graph, clinical codes, arangodb, neo4j
- **Industries:** media, editorial, news, publishing, finance, venture capital, vc, private equity, healthcare, medical, biomedical, clinical, real estate, energy, climate, agriculture, research, b2b sales, retail, home improvement, telecom, historical records, legal
- **Tools:** elasticsearch, qdrant, pinecone, weaviate, cosmos db, chromadb, faiss, neo4j, arangodb, networkx, azure openai, gpt-4, gpt-4o, gpt-5, gemini, llama, langchain, langgraph, autogen, langroid, fastapi, django, streamlit, azure functions, aws lambda, azure ai foundry, mcp, copilot studio, scrapingbee, playwright, puppeteer, fastapi, sentry, docker, gcp, azure
- **Engagements:** Newsweek (Senior AI Engineer, 5 years), CGIAR (agricultural research, 9k+ records), A wealth-management fintech (financial chatbot, 4 agent frameworks), A clinical-intelligence company (20 medical AI agents, 100k+ clinical codes), A private-markets data platform (financial OCR across major VC and PE firms), A news-credibility platform, A market-intelligence startup (100k+ keyword industry analysis), A construction-commerce marketplace (66GB retail catalogs), An HR-tech company (B2B2C)

### Shashwat Verma · Principal · GenAI & Financial Risk

- **URL:** `https://phavella.com/bench/shashwat`
- **Headline:** GenAI for risk + policy design. $150M risk portfolio at Drip Capital, $1B annual transaction policy designed, 16M Citi fraud decisions/month.
- **Skills:** risk modelling, risk strategy, fraud detection, credit risk, underwriting, kyc, aml, neural networks, fraud models, line decrease models, bureau triggers, fair lending, disparate impact, model validation, ks, psi, rob, mrm, model risk management, policy design, transaction policy, governance, audit, multi-agent, agent orchestration, multimodal, document extraction, invoice extraction, prompt engineering, llm agents, prompt workflows, production decision systems
- **Industries:** financial services, banking, lending, trade finance, retail banking, credit cards, invoice finance, venture-backed startups, consumer credit
- **Tools:** python, sql, docker, git, multimodal models, pandas, sklearn, neural networks
- **Engagements:** Drip Capital (Senior Risk Lead, 2 yrs), Citibank N.A. (5+ yrs Risk, retail co-brand card portfolios), A generative-AI automation startup (founder)

## Reference builds · full spines

Four systems, told properly: what broke, what we built, what changed. Every number traces to shipped work.

### Fraud decisioning at scale · Citibank

- **URL:** `https://phavella.com/work/citibank-fraud`
- **Industry:** Banking
- **Headline metric:** 16M decisions / month (fraud decisions / month)
- **Problem:** 16 million card transactions a month need a fraud call in milliseconds. Block too much and you lose customers; too little and you eat the loss.
- **Constraint:** Bank-grade model governance: every model validated (KS, PSI), fair-lending reviewed, and auditable.
- **The system:** Fraud decisioning models across major retail co-brand card portfolios, with policy thresholds tuned per portfolio and a validation trail regulators accept.
- **The human in the loop:** Edge cases route to analysts with full context; the system learns from their calls.
- **Outcome:** 16M decisions/month, 5+ years in production.
- **Outcome figures:**
  - 16M · fraud decisions per month (CITIBANK · 5+ YEARS)
  - 3 · retail card portfolios served (RETAIL CO-BRAND PORTFOLIOS)
  - 5+ yrs · running in production (CITIBANK)

### Risk portfolio & transaction policy · Drip Capital

- **URL:** `https://phavella.com/work/drip-risk`
- **Industry:** Trade finance
- **Headline metric:** $1B policy · $150M portfolio (annual transaction policy · $150M portfolio)
- **Problem:** A fast-growing trade-finance lender extends credit against invoices it has never seen before. Approve the wrong exporter and the loss lands on the book; approve too few and the business stalls.
- **Constraint:** Every decision has to scale with volume without loosening: the policy that governs a $1B flow of transactions must hold as the portfolio grows past $150M.
- **The system:** Risk models and an underwriting strategy feeding a written transaction policy: thresholds, exposure limits, and escalation rules that codify who gets credit, how much, and on what terms.
- **The human in the loop:** Exposures above policy limits route to a risk lead before funds move; the policy captures every decision so the next one is faster.
- **Outcome:** A $150M risk portfolio managed and the policy governing $1B in annual transactions designed.
- **Outcome figures:**
  - $150M · risk portfolio managed (DRIP CAPITAL)
  - $1B · annual transaction policy designed (DRIP CAPITAL)

### Document intelligence at newsroom scale · Newsweek

- **URL:** `https://phavella.com/work/newsweek-doc-ai`
- **Industry:** Media
- **Headline metric:** 30M-doc search (documents indexed)
- **Problem:** A newsroom sits on tens of millions of documents no journalist can search fast enough. The archive is an asset only if you can find the right record in seconds, not hours.
- **Constraint:** Editorial work runs every day; the search and the apps around it have to stay up, stay accurate, and fit into how the newsroom already writes and publishes.
- **The system:** A 30M-document search index behind 15 production AI apps: retrieval, classification, and editorial automation wired into the daily workflow instead of bolted beside it.
- **The human in the loop:** Editors keep the byline: the apps surface, draft, and route; a person approves before anything publishes.
- **Outcome:** 30M-document search, 15 production AI apps, and roughly 497 hours saved every 30 days.
- **Outcome figures:**
  - 30M · documents in one search index (NEWSWEEK)
  - 15 · production AI apps shipped (NEWSWEEK · 5 YEARS)
  - ~497 hrs · saved per 30 days (NEWSWEEK DOC-AI)

### AI video ad production pipeline · Ad Platform (in-house)

- **URL:** `https://phavella.com/work/ad-production-platform`
- **Industry:** DTC & commerce
- **Headline metric:** 5× weekly ad output (weekly ad output)
- **Problem:** DTC brands need fresh video ads faster than a human team can script, shoot, and cut them. Creative volume is the bottleneck on every ad account.
- **Constraint:** Volume can't come at the cost of the brand: nothing runs on an ad account until a person has seen it, so the automation has to leave a clean approval gate.
- **The system:** A 13-state production pipeline that scripts, generates across 11 video models, and voices ads end to end, with human approval gates before anything reaches Meta Ads.
- **The human in the loop:** A human signs off before any ad launches; the pipeline handles the 13 states around that decision, not the decision itself.
- **Outcome:** A pipeline that lifted weekly ad output 5×, with a human signing off before anything launches.
- **Outcome figures:**
  - 5× · weekly ad output (IN-HOUSE AD PLATFORM)
  - 13 · pipeline states orchestrated (IN-HOUSE AD PLATFORM)
  - 11 · video generation models (IN-HOUSE AD PLATFORM)

Also shipped by our principals: Kronus, Creative OS, a clinical-intelligence company, a wealth-management fintech, CGIAR, a construction-commerce marketplace, a private-markets data platform, a life-sciences commercial-analytics firm, a major airline, a North American private-equity firm, a mid-market private-equity firm, a Fortune-tier technology company.

## How we engage · in full

Scope. Build. Ship. Operate. A named process with a human on the calls that matter. Details: `https://phavella.com/engagement`.

### Phases

1. **Scope**: One call, then a written plan: the use case, the data, the integration points, the number it has to move.
2. **Build**: Principals build. Weekly demos on your data, not slides. An eval suite comes with it.
3. **Ship**: Into your stack, your infra, your access controls, with evals and monitoring wired in.
4. **Operate**: We run it with you until it's boring. Then we hand it over, or keep operating it.

Pricing and billing shape are conversation-only, never on public surfaces.

### Boundaries (what we won't do)

Let AI make a consequential decision alone. Rip out your system of record. Free pilots. Equity-only deals. Strategy decks without a build behind them. Engagements where the math doesn't work.

### FAQ

**How does an engagement start?**

With one call. Then a written plan: the use case, the data, the integration points, and the number it has to move. No strategy decks without a build behind them.

**Who actually does the work?**

Four named principals: Parik Ahlawat, Arpit Jhamb, Nikhil Bery, and Shashwat Verma. Principals build; no juniors on production systems. You see weekly demos on your data, not slides.

**Do we have to replace our existing systems?**

No. We layer; we don't migrate. We build from scratch or integrate into the stack you already run (CRM, inbox, docs, ads, ops), and we won't rip out your system of record.

**Does the AI make decisions on its own?**

Routine work auto-executes. Decisions that are expensive to get wrong route to a named human reviewer with context attached and an audit trail. We won't let AI make a consequential decision alone.

**Do you work with startups and small teams?**

Yes. Use cases big and small: from a single workflow to a full platform, idea-stage to in-production. One form takes all requests, and a principal reads every brief.

**What do we get at the end of each phase?**

Scope delivers a written plan, an integration map, and a success metric. Build delivers weekly demos and an eval suite. Ship deploys into your stack with monitoring and runbooks. Operate covers SLAs, tuning, and handover.

**What happens after launch?**

Ship isn't the end. We run the system with you until it's boring. Then we hand it over, or keep operating it.

**What does it cost?**

Pricing is scoped per engagement on the first call. There's no pricing wall on the site. Two boundaries hold either way: no free pilots, and no equity-only deals.

## Thinking · full essays

Notes from building AI that has to work. `https://phavella.com/thinking`

### Integrate, don't replace

- **URL:** `https://phavella.com/thinking/integrate-dont-replace`
- **Published:** 2026-02-18
- **Dek:** The fastest way to kill an AI project is to make it wait for a migration. Layer on top of the stack that already runs.
- **Tags:** integration, system of record, migration, governance

Every operation already has a system of record. It is usually old, partially documented, and load-bearing in ways nobody fully remembers. The instinct, especially among vendors, is to treat that system as the obstacle: migrate first, modernise first, then do the AI. We have watched that sequencing kill more AI initiatives than any model limitation ever has. The migration becomes the project. The AI becomes a line item in year two of a plan that does not survive year one.

The alternative is to treat the existing stack as the substrate, not the obstacle. AI systems are unusually good at sitting on top of things: they read from the CRM without replacing it, draft into the inbox without owning it, reconcile against the ERP without touching its schema. The integration surface (APIs, webhooks, exports, even structured email) is almost always richer than teams assume. When we scoped the Meta Ads integration for our own ad platform, the answer was not a new ads manager; it was a tool server exposing thirty-four operations over the API the team already used. The workflow changed. The system of record did not.

Working this way changes what an integration is. It stops being plumbing and becomes the product decision: which systems does the AI read from, which does it write to, and which does it only ever draft into with a human pressing send. In our video pipeline, the connective tissue is exactly this: Slack for approvals across seven channels, Drive and Sheets and Docs for the assets and the records, the ad platform for delivery. None of those tools were replaced. All of them became more valuable, because the work now flows through them instead of around them.

There is a governance dividend too. When the AI layers over the stack, every action it takes lands in systems the organisation already audits, already backs up, already controls access to. Rip-and-replace resets all of that to zero and asks the company to re-trust a new vendor with everything at once. Layering asks for trust in increments, which is the only way operational trust is actually granted.

The rule we hold: the AI adapts to the operation, not the other way round. If a proposal starts with a migration, it is not an AI proposal: it is an infrastructure proposal wearing a costume. There are moments when replacing a system is right. But that decision should be made on its own merits, on its own timeline, and never as the toll a company has to pay before the useful work is allowed to start.

### What 1,322 automations taught us about 30 industries

- **URL:** `https://phavella.com/thinking/what-1322-automations-taught-us`
- **Published:** 2026-03-05
- **Dek:** The shapes repeat. The last mile never does.
- **Tags:** automations, industries, patterns, exceptions

We keep a dataset of the automation work our principals have shipped and scoped: one thousand three hundred and twenty-two automations across thirty industries, mapped to the tools they run on. We keep it because pattern-matching against real work beats intuition, and because it forces honesty about what actually gets built versus what gets talked about. After enough entries, two things become clear at the same time, and they pull in opposite directions.

The first: the shapes repeat. Strip the vocabulary away and an enormous share of operational AI is one of a handful of moves. Something arrives: an email, a document, a transaction, a call. It gets read and structured. A decision is made about it. It gets routed to a system or a person. Something is written back, and someone is told. Intake, extraction, decision, routing, record. A logistics dispatcher and a clinic front desk do not think they share a workflow, but at this altitude they nearly do. This is why experience compounds across industries: the engineer who has shipped document extraction for private equity fund reports has already solved most of the retail catalogue problem, even though the two clients would never attend the same conference.

The second: the last mile never repeats. The recurring shape is maybe seventy percent of the system, and it is the easy seventy percent. What remains is the part that decides whether the thing survives: this company's exception rules, the field that means something different in this ERP than in every other ERP on earth, the approval that legally must be a named human in this jurisdiction, the supplier who still sends the manifest as a photographed fax. None of that generalises. All of it is the actual work. This is why we refuse to group industries into neat clusters and sell the cluster: the label flatters the pattern and hides the last mile, and the last mile is where projects die.

The practical consequence is a two-speed way of building. Move fast through the recurring shape. It is known territory, and there is no prize for rediscovering it slowly. Then slow down, deliberately, at the boundary where the operation becomes itself. Walk the exceptions with the people who handle them today. Read the weird records, not the clean ones. Budget real time for the fax.

When someone asks whether we know their industry, the honest answer is usually: we have shipped the shape of your problem more than once, and we have never shipped your version of it. Both halves of that sentence matter. The first is why the project will not start from zero. The second is why we will not pretend it starts at ninety percent.

### Bank-grade is a discipline, not a badge

- **URL:** `https://phavella.com/thinking/bank-grade-is-a-discipline`
- **Published:** 2026-03-24
- **Dek:** Sixteen million fraud decisions a month will teach you what validation actually means.
- **Tags:** financial services, model validation, fair lending, risk policy

Every vendor deck in financial services says bank-grade somewhere. Almost none of them can describe the homework. Bank-grade is not an adjective. It is a set of obligations: every model validated before it touches a decision, every threshold documented with the reasoning attached, every decision reconstructable months later when someone asks why. One of our principals spent more than five years inside that regime, running fraud decisioning that makes sixteen million calls a month. What follows is the unglamorous core of it.

First obligation: prove the model discriminates, in the statistical sense, and only that sense. Validation metrics like KS tell you whether the model actually separates good transactions from bad ones; stability metrics like PSI tell you whether the population it sees in production still resembles the population it was built on. These are not academic niceties. A model that silently drifts as the customer base shifts is not a model, it is a liability with a dashboard. The discipline is running these checks on a schedule and treating a breach as an incident, not a curiosity.

Second obligation: fairness review is part of the definition of working. A fraud or credit decision system that performs brilliantly on aggregate and unevenly across protected groups is a failed system, full stop. Regulators will say so eventually, but the point is that it is true before they say so. Designing for fair-lending review from the start changes feature selection, changes documentation, and changes who signs off. Retrofitting it after launch is somewhere between painful and impossible.

Third obligation: the policy layer is a first-class artifact. The model produces a score; the policy decides what happens at that score, per portfolio, per product, with the trade-offs written down. When our principal designed the risk policy governing a billion dollars a year in transactions at a trade-finance lender (alongside a hundred-and-fifty-million-dollar portfolio) the policy document was the deliverable. The model was an input to it. Teams that ship the score without the policy have shipped half a system and kept the dangerous half.

None of this is specific to banks anymore. The moment your AI touches money, eligibility, or anything a customer can be wrongly denied, you have inherited the obligations whether or not a regulator is watching yet. The good news is that the discipline is learnable and mostly procedural. The bad news is that there is no shortcut through it, and any system sold to you without validation, stability monitoring, fairness review, and a written policy layer is asking you to carry the risk it did not bother to engineer away.

### Retrieval at 30 million documents

- **URL:** `https://phavella.com/thinking/retrieval-at-30-million-documents`
- **Published:** 2026-04-14
- **Dek:** Search quality is an evaluation problem long before it is an infrastructure problem.
- **Tags:** retrieval, enterprise search, OCR, evaluation

A search index with thirty million documents in it is not a bigger version of a search index with thirty thousand. Somewhere between those two numbers, every weak assumption you made becomes load-bearing and then becomes visible. One of our principals spent five years building this at a national newsroom: the archive, the extraction pipelines, and the fifteen production AI applications that sat on top of it. The lesson that survived all five years: retrieval quality is an evaluation problem long before it is an infrastructure problem.

Teams reach for infrastructure first because infrastructure is legible: swap the index, add a vector store, tune the analyzer. But none of those moves can be judged without a fixed set of real queries and judged answers. Build that eval set first, from the questions people actually ask, including the ugly ones with typos and half-remembered names. Fifty honest queries with graded results will steer an architecture better than any benchmark blog post, because they encode your corpus, your users, and your definition of correct.

The corpus itself is where most of the quality lives. At newsroom scale, extraction is not a preprocessing step. It is the product's foundation. Optical character recognition on decades of scans, deduplication across syndicated copies, entity normalisation so a person's four name variants resolve to one. We have hit the same wall in every domain that keeps records: sixty-six gigabytes of retail catalogues, fund documents in private equity, nine-thousand-plus records in an agricultural research archive. Different industries, same truth: if the extraction is sloppy, no retriever downstream can be smart enough to compensate.

What made the newsroom work count was not the index. It was that editors stopped spending their mornings hunting for context. Measured against the prior workflow, the document systems were saving on the order of five hundred hours in a thirty-day window. That number is the point. Retrieval systems justify themselves in recovered hours and faster decisions, and if you cannot trace the connection from relevance scores to someone's Tuesday, the system is not done.

So the order of operations we hold: eval set, then extraction, then index, then model. Most teams run it exactly backwards. They pick the most interesting technology first and discover at the end that nobody can say whether the answers got better. At three hundred documents you can get away with that. At thirty million, the corpus will collect the debt.

### In regulated care, design the escalation first

- **URL:** `https://phavella.com/thinking/healthcare-ai-escalation-first`
- **Published:** 2026-05-06
- **Dek:** Twenty medical agents in production: the design that made them safe to run.
- **Tags:** healthcare, medical agents, escalation, guardrails

The interesting thing about putting twenty medical AI agents into production is not the agents. It is the order in which they had to be designed. In most domains you can build the happy path first and harden the edges later. In healthcare that sequence is backwards to the point of negligence: the escalation path (what the system does when it should not act alone) is the first thing to design, because it is the thing that makes everything else permissible.

Escalation-first changes the architecture in concrete ways. Every agent needs a defined refusal boundary: the set of situations where the correct output is not an answer but a handoff, with the context packaged so the human picks up mid-stride rather than starting over. The handoff is not a failure mode to minimise. It is a designed outcome to get right. The systems that scare people are not the ones that escalate often; they are the ones that never do.

The second requirement of the domain is speaking its language, literally. Clinical work runs on controlled vocabularies, and grounding the agents in more than a hundred thousand clinical codes was not optional plumbing. It was the difference between an assistant that gestures at medicine and one that can participate in it. Every regulated field has an equivalent spine: the terminology, the codes, the canonical identifiers. Ground in them early. Free-text cleverness does not survive contact with a compliance review.

We saw the same pattern hold in finance when one of our principals built a client-facing chatbot for a wealth platform, orchestrating four different agent frameworks under one set of guardrails. Different regulator, same architecture: hard boundaries on what the system may claim, glossaries it must respect, an escalation surface a human actually staffs. Regulated domains do not reward the boldest system. They reward the system whose limits were designed with the same care as its capabilities, and that turns out to be the system that ships.

### Volume is a system property

- **URL:** `https://phavella.com/thinking/volume-is-a-system-property`
- **Published:** 2026-05-27
- **Dek:** Five times the ad output did not come from better prompts. It came from a pipeline with thirteen states.
- **Tags:** creative automation, pipelines, video generation, approvals

When our video ad platform reached five times the weekly creative output, nobody's prompts had gotten five times better. Prompts are not where volume comes from. Volume is a system property: it falls out of how work moves between states, where humans sit in the flow, and what happens automatically while they sleep. The generative model is the loudest component and close to the least important one.

The platform runs production as a thirteen-state pipeline, and the states are the design. A brief does not become a published ad in one heroic generation. It moves: scripted, generated, voiced, assembled, reviewed, revised, approved, delivered, measured. Each transition has an owner, and the owner is only sometimes a person. The discipline of naming the states did more for throughput than any model upgrade, because it turned an artisanal process into an operable one. You cannot parallelise what you have not named.

Creative systems also refuse to depend on a single model gracefully. The pipeline orchestrates eleven video generation models, not because more is better, but because each has a personality: one is right for product shots, another for motion, another for people. Routing between them is a production decision the system makes per brief. Model plurality is what turns 'the model is having a bad day' from an outage into a routing event.

And the human gates are why the volume is usable. Five times the output of unreviewed creative is a liability multiplier, not a growth story. Approvals run through the channels the team already lives in (review in Slack, assets in Drive, records in Sheets) so brand judgment stays human while everything around the judgment is automated. That is the general recipe, and it has nothing specific to advertising in it: name the states, route around individual model weakness, and put the human sign-off exactly where the risk is. Scale the system around the judgment, never in place of it.

### Why most AI never reaches production

- **URL:** `https://phavella.com/thinking/why-most-ai-never-ships`
- **Published:** 2026-06-01
- **Dek:** The gap between a demo and a system that runs is where most AI projects die. Here's what closes it.
- **Tags:** production AI, evaluation, data quality, edge cases

A demo is a proof of possibility. A production system is a proof of reliability. The distance between them is not measured in engineering hours. It is measured in how honestly a team has confronted their data, their edge cases, and the cost of being wrong. Most AI projects die in that gap, not because the model failed, but because the surrounding architecture was never designed to carry real load.

The first thing that collapses in production is the assumption about data quality. In a demo, you choose your inputs. In the field, you get whatever arrives: malformed records, missing fields, schema drift between the CRM and the data warehouse, documents scanned at 72 dpi by someone in a hurry. A model that performs at 94% on your curated eval set can drop to something genuinely embarrassing on the first month of live traffic. The engineers who built the demo did not lie; they just never had to face real data.

The second thing that collapses is evaluation. Most teams ship without a measurement framework and then discover they have no idea whether the system is working. They rely on the absence of complaints, which is not a signal. It is a silence. Good production AI has a defined eval set, a set of failure modes that are treated as first-class bugs, and a process for reviewing outputs before and after model changes. This is not glamorous work. It is the work that separates a system from a science project.

Third: operating an AI system is not the same as deploying one. Costs drift upward as usage grows. Models update and behavior shifts. A new edge case surfaces at 2am. The teams that ship and sustain production AI treat these as normal operational concerns from day one: they instrument everything, set cost budgets, and build human review into the loop for anything consequential. The teams that do not treat these as first-class concerns end up firefighting them six months in, usually right before a contract renewal.

The path forward is not to de-risk the model. It is to de-risk the system. That means starting smaller than you think: a bounded workflow with a clear success criterion, not a platform. It means investing in data before you invest in architecture. It means designing the evaluation harness in week one, not week twelve. And it means being honest about what the business needs versus what makes a good slide. Most of the AI that actually runs in production is not the most technically impressive version of what was possible. It is the version that was honest about its constraints and built accordingly.

### Human-in-the-loop is a feature, not a disclaimer

- **URL:** `https://phavella.com/thinking/human-in-the-loop-is-a-feature`
- **Published:** 2026-06-10
- **Dek:** The teams that win with AI design the sign-off in, not bolt it on.
- **Tags:** human-in-the-loop, review, consequential decisions, training data

There is a version of human-in-the-loop that is apologetic: a hedge, a liability clause, a way of saying we do not fully trust what we built. That version shows up as a review step that no one has time for, staffed by someone who was not told what to look for, with no record of what they approved or why. That version is not a feature. It is theatre.

The version that actually works starts from a different premise: that most decisions in any operation are routine, and a small number are consequential. Routine decisions (categorising a transaction, routing a support ticket, flagging a document for review) can and should auto-execute. Consequential decisions (approving a contract, declining a loan, publishing a customer-facing communication) need a named person, with context, and a record. The design question is not whether to have humans in the loop. It is where to put them.

When you get this right, the AI does not slow the organisation down. It frees the organisation to think clearly about the decisions that matter. The reviewer is not reading a wall of raw output; they are reading a structured briefing that the system prepared, with the relevant evidence surfaced and the recommendation stated plainly. Their job is to confirm or override, and either action is logged. Six months later, when something goes wrong, you know exactly who saw what and what they decided. That is not bureaucracy. That is how consequential decisions should work.

The operational benefit compounds. The review log becomes training data. Patterns in overrides reveal where the model is consistently wrong or where the policy is unclear. Teams that build this way learn faster than teams that ship a black box and wait for complaints. They also build trust with the people who use the system, because those people can see the mechanism and know there is a hand on the wheel.

Our own systems are built this way from the start. The approval surface is not an afterthought. It is a first-class piece of the design, alongside the model, the eval set, and the data pipeline. When we propose a system, we specify which decisions auto-execute, which decisions route to a named reviewer, what context the reviewer gets, and how the audit trail is structured. That specification is as important as the model choice. In some engagements it is more important.

### Strategy that ships

- **URL:** `https://phavella.com/thinking/strategy-that-ships`
- **Published:** 2026-06-20
- **Dek:** AI strategy is worthless as a deck. It is worth a lot as a working proof.
- **Tags:** strategy, discovery, proof, risk management

We have seen the deck. Most organisations that come to us have already paid for the deck. It describes a transformation, a roadmap, a set of use cases ranked by impact and feasibility, a reference architecture, and a note about change management. It is usually well-written. It is usually not acted on. Not because the thinking was wrong, but because a deck cannot tell you what you will actually find when you connect to the real data and run the first model against the first workflow.

The alternative is discovery that writes code. Not a blueprint for a future system. A contained, working proof against a real use case, built in a fixed time window, against the actual data, with a defined success criterion that both sides agree on before work starts. The output is not a recommendation. It is a system, or the first piece of one, that the organisation can observe running. The strategy emerges from what you learn.

This matters for risk management as much as it matters for learning. A bounded diagnostic with a fixed fee is a different risk profile than an open-ended transformation engagement. The client knows what they are paying, what they will receive, and what question will be answered. We know what we are accountable for. Neither side is betting on a plan that has not touched reality yet.

What we build in a discovery is explicitly designed to scale into production or to be thrown away, and both outcomes are fine. If the proof works, the next phase has a foundation: real data, real eval results, real edge cases documented. If it does not work, the client has spent a defined budget to learn something true, which is more valuable than spending a larger budget to learn something wrong six months later. The honest version of this conversation is that we cannot always tell you in advance which outcome you will get. We can tell you that you will know sooner, and that knowing sooner is worth paying for.

### What AI can actually automate in a 3PL back office (and what it can't)

- **URL:** `https://phavella.com/thinking/automating-the-3pl-back-office`
- **Published:** 2026-07-05
- **Dek:** A 3PL back office runs on rekeying. AI can take the typing, the anomaly-spotting, and the status-chasing. It should not touch the rate exceptions or the disputes.
- **Tags:** logistics, automation, back office, human-in-the-loop

The back office of a third-party logistics operation is, at its heart, a rekeying machine. An order arrives by email. Someone reads it and types it into the transportation management system. The same numbers get typed again into the accounting package, and often a third time into a customer portal or the spreadsheet a manager actually trusts. Nobody chose this arrangement. It accreted, one integration gap at a time, until the fastest way to move data between two systems became a person with a keyboard. This is the work AI is genuinely good at removing, and it is worth being precise about which parts of it, because the precise version is more useful than the excited one.

Start with what automates cleanly. Reading a document and entering it somewhere is the core competence of current models. A rate confirmation, a bill of lading, a proof of delivery, a carrier invoice: these are semi-structured documents that arrive in predictable-enough shapes, and a model can pull the fields, match them against an existing load, and stage the entry into the transportation system. Flagging anomalies is the second thing that automates well: an invoice that does not match the quoted rate, a shipment whose weight jumped between booking and manifest, a delivery that has gone quiet past its window. Chasing status is the third: the endless outbound of 'where is my truck' messages that a system can send, read the replies to, and summarise, without a person drafting each one by hand. When the match against the existing load is clean, that entry can move on its own; when a field is missing or the numbers disagree, it waits for a human glance. Between them, these three account for most of the hours a back office actually burns, which is why taking them off people's desks is worth more than the phrase makes it sound.

What those three have in common is that they are reversible and legible. If the model stages the wrong field, a person catches it at the entry it prepared, before anything downstream moves. If it flags an anomaly that turns out to be nothing, the cost is a few seconds of a human glance. If it chases a status that was already resolved, a driver is mildly annoyed. None of these failures compound. That is the real test for what routine work can safely auto-execute: not whether the model is clever, but whether being wrong is cheap and caught early.

Now the other side of the line, which is the more important half. A 3PL back office is full of decisions that look like data entry and are actually judgment. A rate exception, where a customer wants a price the standard tariff does not give them, is a commercial call with a relationship behind it. A carrier dispute, where the invoice is over and the carrier insists the detention was real, is a negotiation, sometimes with a contract clause in play and always with a working relationship you would rather not spend. Whether to hold a shipment, waive a fee, or eat a cost to keep an account happy: these are the calls that decide whether the operation makes money and keeps its customers. They are exactly the calls a model should not make alone. Not because it cannot produce a plausible answer, but because a plausible answer to a consequential question, executed without a human, is how you lose an account quietly and find out at renewal.

So the design is not 'automate the back office.' It is 'let the routine flow through the system and route the consequential to a person, with the context already assembled.' The model reads the carrier invoice, matches it to the load, and when the numbers agree it stages the entry and moves on. When they do not, it does not decide the dispute. It packages the discrepancy, the quote, the invoice, and the relevant emails, and puts it in front of the person who owns that carrier relationship, who now spends half a minute deciding instead of ten minutes assembling. The human is not removed from the loop. The human is moved to the part of the loop that needed a human all along, and freed from the part that never did.

The reason this is buildable at all is that we do not replace the transportation system or the accounting package to do it. We layer on top of them. Those systems already expose the surface an AI needs: an API, an export, a structured inbox, sometimes just a webhook, and where none of those exist a structured email is often enough to read from and write back to. When we built the integration for our own ad platform, the answer to connecting the ad account was not a new ads manager; it was a tool server exposing thirty-four operations over the interface the team already used. A 3PL back office is the same shape of problem. The transportation system stays the system of record. The accounting package stays the book. The AI reads from one, drafts into the other, and asks a human before it does anything that cannot be quietly undone. There is a quieter dividend in this arrangement, too: because the AI acts through systems the operation already audits and backs up, every entry it stages lands somewhere the company can already see, rather than in a new tool nobody yet trusts. The integration is the product decision, and it is a decision to fit the operation rather than rebuild it.

The honest version of the pitch is unglamorous, and better for being so. We are not going to hand anyone an autonomous back office that runs itself while the team sleeps. We are going to take the rekeying, the anomaly-spotting, and the status-chasing off people's plates, and we are going to leave the rate exceptions, the disputes, and the account-shaping judgment exactly where they are, with the people whose job that is, now moving faster because the context arrives assembled. The back office gets quieter. The decisions that matter stay human. That is the trade we would put our name to.

### Document intelligence for clinic groups: intake without the retyping

- **URL:** `https://phavella.com/thinking/document-intelligence-for-clinics`
- **Published:** 2026-07-06
- **Dek:** Multi-location clinics type the same patient details three times. Document intelligence reads the referral and drafts the entry; a person approves in one click, and nothing patient-facing runs alone.
- **Tags:** healthcare, document intelligence, intake, retrieval

A clinic group's front desk is where the same patient information gets typed into existence three or four times. It arrives on a paper intake form, or a portal the patient half-filled. A staff member keys it into the practice management system. When a claim goes out, the relevant subset is entered again into the billing flow, and if a referral came in, someone has already read a PDF from another practice and transcribed the parts that matter. Multiply that by every new patient across every location and you have a standing tax on the people you least want doing data entry: the ones who are supposed to be looking after patients. Each retyping is also a fresh chance to mistype a date of birth or an insurance number, and each of those errors is one a person downstream has to notice, chase, and unwind. The tax is not only the minutes at the keyboard; it is the slow accumulation of small mistakes that someone, eventually, has to pay for.

The referral pile is the sharpest version of the problem. A referral arrives as a PDF, often a fax that became a PDF, from a system that does not talk to yours. It carries a name, a date of birth, an insurance identifier, a reason for referral, sometimes a medication list and a page of history. Every field on it already exists somewhere in a computer. It just exists in the wrong computer, in a format your practice management system cannot ingest, so a human reads it and retypes it. This is precisely the shape of work document intelligence removes: read the document, pull the fields, draft the entry. It is worth being clear that the model is not being asked to understand medicine here. It is being asked to move known fields from a page into a form, a narrower and far more reliable task than the word intelligence sometimes implies.

The design that makes this safe in a clinical setting is that the system drafts and a person approves. The model reads the referral, extracts the fields, and stages a practice-management entry with each value it filled and the place on the page it took it from. A staff member sees the draft next to the source and confirms it, or corrects the one field that was ambiguous, in a single pass. Nothing about a patient is written unattended. The point is not that the model is trusted; it is that the model is checked, cheaply, at the moment of entry, by the person who would have been typing the whole thing anyway. Their job changes from transcription to review, which is faster and less error-prone than typing from a fax. It changes the failure mode, too. A mistake now has to survive a person looking straight at the source, rather than slip through unseen because nobody had time to double-check the typing. In a clinical setting the location of the rare error is exactly what you are designing around, and this puts it in the safest place there is: in front of a human, at the one moment the source and the entry sit side by side.

People reasonably ask whether this holds up beyond one referral at a time, and the honest answer comes from scale we have actually run. One of our principals spent five years building document intelligence at a national newsroom: a search index of thirty million documents, fifteen production AI applications sitting on top of it, and, measured against the prior workflow, recovery on the order of close to five hundred hours in every thirty-day window (the exact figure was around four hundred and ninety-seven). A clinic group is a smaller corpus than a newsroom archive, not a different kind of problem. The mechanics are identical: extract cleanly, index so it is findable, and keep a human on anything that gets acted on. If the pattern holds at thirty million documents, it holds at a clinic's intake volume with room to spare. The same reading that drafts the entry also makes the referral findable afterward, which is the second and quieter payoff: instead of a folder of PDFs nobody can search, the clinic ends up with records a clinician can pull up in seconds. Intake is simply where the retyping hurts most, so it is the right place to begin.

The part that does not transfer by analogy is the compliance envelope, and that has to be built as a requirement rather than bolted on afterward. Handling patient information means the constraints, who may see what, where it is stored, what is logged, what is allowed to leave the building, are inputs to the design from the first day, not a review the system tries to pass at the end. We will say plainly what that does and does not mean. It means the human sign-off on anything patient-facing is not optional and not negotiable. It means the record of who approved which entry is a first-class part of the system. It means the data can be kept where the regulation requires it to stay. It does not mean we hand you a certification, and we will never claim one we do not hold. Compliance is a discipline the whole system is built to respect, not a badge to wave.

What the clinic actually gets is quieter than the pitch usually sounds, and better for it. The referral pile stops being a queue of retyping and becomes a queue of one-click confirmations. New-patient intake stops being typed three times and is entered once, from the source, with a person checking rather than transcribing. The staff who were doing the keying are doing the reviewing, which is the part that needed a human, and they are doing less of it per patient. Nothing patient-facing runs on its own. The system does the reading and the drafting; the clinic keeps the sign-off. That division, machine reads and human approves, is the whole design, and it is the same one we would build for a newsroom, a lender, or a logistics desk, moved into a setting where the stakes make the human gate non-negotiable.

### The AI app you ship in eight weeks: what production-ready means

- **URL:** `https://phavella.com/thinking/what-production-ready-means`
- **Published:** 2026-07-07
- **Dek:** A demo is a proof of possibility. Production-ready is a proof of reliability, and it means a specific set of things or it means nothing.
- **Tags:** production, evaluation, shipping, AI apps

'Production-ready' is the most quietly abused phrase in AI work. It gets attached to a demo that worked once, on a laptop, on data someone chose. A demo is a proof that the thing is possible. Production-ready is a proof that it is reliable, and the distance between those two claims is where most AI projects quietly die. It is worth saying concretely what closes that distance, because production-ready should mean a specific set of things or it means nothing at all.

A demo and a production system can look identical in a screen recording. They differ in what you cannot see. The demo runs on inputs that were, consciously or not, curated: the clean record, the well-lit invoice, the question phrased the way the builder expected. It has no evaluation set, so nobody can say whether the next answer will be as good as the one on screen. It has no monitoring, so when it drifts, the first signal is a complaint. It has no error path, so the moment a malformed input arrives, the behaviour is undefined. And it has no human lane for the cases that should not be decided by a model. None of that shows up in the recording, which is exactly why the demo is convincing and the production system is hard. The demo is not dishonest about any of this. The people who built it never had to meet the inputs they did not choose, and everything that makes production hard lives precisely in the inputs nobody would put on a slide: the malformed, the empty, the surprising, the ones that arrive at the worst possible moment.

So the first artifact we write is the evaluation set, before the model, before the interface, before anything that photographs well. It is a fixed collection of real inputs with judged correct answers, drawn from the actual data including its ugly parts: the typo, the half-missing record, the request phrased three different ways. A threshold gets attached to it, the level of performance below which the system does not ship. This is unglamorous, and it is the single highest-leverage decision in the build, because everything downstream (which model, which retrieval, whether a given change helped or hurt) can only be judged against it. Teams that skip it are not moving faster. They have just chosen not to know whether the thing works.

The second thing production-ready means is that the boring ninety percent is handled. A demo shows the interesting path: the clever extraction, the fluent answer. A production system spends most of its code on the unglamorous remainder. What happens when the input is empty, malformed, or in a format nobody anticipated. What happens when the model is slow or unavailable. What happens when confidence is low. What gets logged, what gets retried, what gets escalated. This is ordinary software engineering, it is most of the work, and its absence is the difference between something that survives a Tuesday and something that only survives a rehearsal. Teams underweight this remainder consistently, because it is invisible in every artifact that ever gets shown: the slide, the recording, the update to the sponsor. It becomes visible in exactly one place, the first week of real traffic, which is the most expensive place there is to discover it.

Third, there is a human review lane, and it is designed rather than improvised. Routine outputs auto-execute; the consequential ones route to a named person with the context assembled and their decision logged. This is not a hedge against a weak model. It is how a serious system draws the line between what is cheap to get wrong and what is not, and it is the same design whether the domain is a newsroom, a lender, or a clinic. Six months later, when someone asks who approved a particular output and why, the system can answer. That answer is part of what production means.

This is why the engagement has the shape it does. Scope produces a written plan and a single success metric both sides agree on before any code is written. Build runs as weekly demos on your data, not slides, so the thing is exercised against reality from the first week rather than the last, and an eval suite comes with it. Ship deploys into your stack, behind your systems, wired to the tools you already run, with monitoring and runbooks. Operate means we stay until it is boring: the monitoring quiet, the escalations rare, the surprises gone. A bounded first build of this shape tends to fit in roughly eight weeks, not because eight weeks is magic, but because a scope that cannot reach boring-in-production in that window is usually a scope trying to do too much at once. The weekly rhythm is the real safeguard. A system exercised against your data every week cannot hide its failures until the end, because there is no end to hide them in, only the next demo. Problems that would otherwise ambush a launch surface in week two instead, when they are cheap to fix and nobody has staked a date on them yet.

The evidence that production is a real and reachable bar, not a marketing word, is that our principals have cleared it repeatedly. One built fifteen production AI applications on top of a thirty-million-document index at a national newsroom. Not fifteen demos: fifteen systems editors depended on every day, with the search staying up and the outputs staying accurate inside the workflow people already used. Fifteen is a number you only reach if production means something specific and you hit it every time. That is the bar we build to. The demo is the easy part, and we treat it as the start of the work, not the end of it. We lead with that record not to impress but to set an expectation. When we say we will make something production-ready, we mean the specific list above, and we would rather lose a deal by being clear about that bar than win one by quietly lowering it.

### From dashboards to decisions: what AI analytics is for

- **URL:** `https://phavella.com/thinking/from-dashboards-to-decisions`
- **Published:** 2026-07-08
- **Dek:** A dashboard shows the past and waits for a human to act. A decision system reads the data, proposes the call, and keeps a person on the ones that matter.
- **Tags:** analytics, decisions, data, policy

Most analytics stops one step short of the point. A dashboard gathers the data, computes the numbers, and arranges them so a human can look. Then it waits. Someone has to notice the chart, interpret it, decide what it means, and act, and that someone is usually busy, so the chart is often noticed late or not at all. The dashboard is not wrong. It is incomplete: it shows the past and leaves the decision, which is the whole reason the data was gathered, entirely to a person who may never get to it. The dashboard's implicit promise was that seeing the number would be enough. In practice, seeing is the cheap half. The expensive half is the judgment and the action that were supposed to follow, and those it quietly hands back to a person and hopes for the best.

The gap has a cost that dashboards hide. A number on a screen is only worth the decision it eventually informs, and between the number and the decision sits human attention, the scarcest resource in any operation. A runway figure that is accurate and unread changes nothing. A spike in a churn cohort that nobody clicks into is not insight; it is a missed alarm with good production values. The dashboard measures the past faithfully and then asks a human to supply the part that actually matters, on their own time, from memory, under load. The chart does the easy part perfectly and leaves the hard part undone, then presents the easy part as if it were the deliverable.

AI analytics is worth building when it closes that loop. Not another chart, but a system that reads the data the dashboard would have shown, works out the decision it implies, and either proposes that decision to a person or, for the routine cases, takes it. The output is not 'here is a number.' It is 'here is the decision, and here is the reason,' with the evidence attached. The shift is small to describe and large in practice: the system stops handing a human a fact to interpret and starts handing them a recommendation to confirm or override. A human is still in the loop for anything consequential. They are simply no longer the component that has to notice.

The clearest version of this we have built is a risk policy, not a dashboard. One of our principals designed the risk system at a trade-finance lender that manages a hundred-and-fifty-million-dollar portfolio and wrote the policy governing a billion dollars a year in transactions. A dashboard there would have shown exposure by exporter and waited. The system instead holds a decision: it scores each invoice against the policy, and where the exposure sits within limits it underwrites, and where it exceeds them it routes to a risk lead before any funds move. The analytics is not reporting on the lending. It is making the lending decision, per transaction, under a written policy, with a human on the calls above the line. That is the difference between analytics that describes and analytics that decides. Note what the system did not do: it did not remove the human. The exposures above the limit are still a person's call, made with the exporter, the amount, and the reason laid out in front of them. What it removed was the lag between a number existing and a decision happening, which at a billion dollars a year in flow is not a small thing to have removed.

The reason a policy is the right artifact, rather than a smarter chart, is that a decision system has to be legible and it has to hold. Legible: every decision it makes carries its reason, so months later someone can reconstruct why a particular invoice was approved or held. Holds: the policy applies the same thresholds at ten transactions and at ten thousand, so growth does not quietly loosen the standard. A dashboard has neither obligation, because a dashboard never decides anything; it is the human downstream who silently supplies the judgment and the consistency, or fails to. Moving the decision into the system is what makes the judgment auditable and the consistency real. It is also what makes the routine safe to leave alone. Because every decision carries its reason, a wrong pattern shows up as a pattern a reviewer can catch, rather than as a scatter of one-off complaints. A dashboard leaves no such trace of the decisions taken off the back of it, so when something drifts there is nothing to audit, only memories to argue over.

This is also where analytics meets the rule we hold everywhere: a human stays on the consequential decisions. Closing the loop does not mean handing the model the keys. The routine calls, the invoice comfortably within policy or the metric moving inside its normal band, can auto-execute, because being wrong there is cheap and caught. The consequential calls, the exposure over the limit or the anomaly that would change a hiring plan, route to a named person with the recommendation and the reason in front of them. The system does the reading, the arithmetic, and the first draft of the judgment. The person does the part that should never be automated, and does it faster because the case arrives assembled. The division of labour is the same one we draw everywhere: the machine handles volume and first drafts, the person handles the calls that are expensive to get wrong, and neither is asked to do the other's job.

So the question to ask of any analytics project is not 'what will this show me' but 'what decision will this close.' If the honest answer is that it will produce a handsome surface someone still has to interpret and act on from memory, it is a dashboard, and it will join the other dashboards nobody opens. If the answer is that it will read the data and put a decision, with its reason and a human on the ones that matter, in front of the person who owns it, that is analytics worth building. The value was never in seeing the number. It was always in the decision the number was for.

### Fraud screening at 16 million decisions a month: what holds

- **URL:** `https://phavella.com/thinking/fraud-screening-what-holds`
- **Published:** 2026-07-09
- **Dek:** Sixteen million decisions a month for more than five years. What keeps a screening system reliable is never the model.
- **Tags:** fraud, decision systems, evaluation, scale

A fraud screening system that makes sixteen million decisions a month, and has made them for more than five years, is a useful thing to study, because at that scale the illusions burn off. One of our principals ran exactly this: card-transaction fraud decisioning across three national retail portfolios, in production, for over five years. The question people expect to be interesting is which model it used. The question that actually determines whether such a system survives is quieter: what holds it up, month after month, when the model is only one component among many. The answer is never the model. It is the discipline around it. That is not modesty about models, which matter. It is an observation about where reliability comes from at scale. Across five years the models were revised more than once; the thing that stayed constant, and kept the system trustworthy through those revisions, was everything wrapped around them.

Start with the thing the model does, so it can be set aside. It produces a score, a number that says how likely this transaction is to be fraud. That is genuinely useful and completely insufficient. A score is not a decision. Sixteen million times a month, something has to turn that number into approve, decline, or hold, and that translation is where a screening system lives or dies. It is done by thresholds, set per portfolio, tuned to the trade-off each business will accept between blocking good customers and eating losses. The model is the sensor. The policy on top of the score is the system, and it is the part that has to be designed, documented, and owned.

The second thing that holds is a review lane for the cases the system should not decide alone. Most transactions are easy: clearly fine and cleanly cleared, or clearly bad and cleanly blocked. The danger lives in the uncertain middle, the transactions the score cannot confidently place. Sending those straight to a decision, in either direction, is how a system either bleeds fraud or infuriates good customers at scale. So the uncertain ones route to analysts, with the context assembled, and a human makes the call. This is not the system failing. It is the system working as designed: the human valve sits exactly on the band where being wrong is expensive, and the analysts' decisions feed back as labels that sharpen the next round. Sizing that lane is its own discipline. Route too little to people and the uncertain cases get decided badly at machine speed; route too much and you have rebuilt the manual process you were trying to escape. The threshold that decides what counts as uncertain is not a detail. It is one of the most consequential numbers in the system, set deliberately, reviewed, and moved only when the evidence says to.

The third thing that holds is the audit trail, and at this scale it is not optional bookkeeping. Every decision has to be reconstructable months later: what the score was, which threshold applied, what the outcome was, and if a human touched it, who and why. Part of that is for regulators, who will eventually ask. The deeper reason is that a decision system you cannot inspect is a decision system you cannot trust, debug, or improve. The trail is what turns sixteen million opaque calls a month into something a team can actually reason about when one of them goes wrong, as some inevitably will. When one does, the trail is the difference between a cause found in an afternoon and a week of guessing. You can replay the exact inputs, see which threshold applied, and tell whether the failure was the model, the policy, or a human override. That is knowledge you do not have if the system only remembers its final answers.

The fourth thing, and the one most teams underweight, is the evaluation that catches drift before customers do. A model that was accurate when it launched is not permanently accurate. The population it screens shifts: new customers, new fraud patterns, new products, a holiday season that looks nothing like the training data. If nobody is watching for that drift, the first signal is a rise in complaints or losses, which means customers absorbed the failure first. The discipline is to monitor the population and the performance on a schedule, and to treat a meaningful shift as an incident to investigate, not a curiosity to note. Catching drift in the evaluation is the difference between a quiet fix and a public one. Customers should never be the drift detector. If the first place a shift becomes visible is the complaint queue, the monitoring was decorative.

We have written elsewhere that bank-grade is a discipline, not a badge, and this is the operational core of that idea. None of the four things that hold a screening system up is a model choice. Thresholds, a human review lane, an audit trail, and drift-catching evaluation are all system properties: designed, staffed, and maintained, and constant while models come and go underneath them. A better model dropped into a system without these buys you a slightly better score and none of the reliability. The same model inside a system that has these holds for five years and sixteen million decisions a month. This is the uncomfortable part for anyone selling a model as the answer. The durable work is not the part that demos well. It is the thresholds argued over in a room, the review lane someone has to staff, the trail nobody looks at until they need it, and the drift checks that run whether or not anything is wrong.

This is why, when the stakes involve money or eligibility, we build the system before we get precious about the model, and why the same shape shows up in adjacent work. The risk policy at a trade-finance lender, governing a billion dollars a year in transactions against a hundred-and-fifty-million-dollar portfolio, holds for exactly the same reasons: score, threshold, a human valve on the uncertain, an audit trail, and a watch for drift. The model is the part everyone wants to talk about and the part that matters least to whether the thing survives contact with real volume. What holds is never the single call the model makes. It is the system that decides what to do with it.

### Why your e-commerce ad pipeline stalls at ten variations

- **URL:** `https://phavella.com/thinking/why-ad-pipelines-stall`
- **Published:** 2026-07-10
- **Dek:** Ad performance runs on variation, and hand-run creative stalls around ten a week. The bottleneck is production, not ideas.
- **Tags:** creative, ads, pipeline, DTC

Direct-response advertising runs on variation. The single best ad is worth less than the ability to keep putting fresh ones in front of an audience that fatigues on everything eventually. Every performance marketer knows this, which is why the brief is always 'more creative,' and every performance marketer also knows the wall it hits: a hand-run creative process tends to top out around ten variations a week, and then it stalls. The instinct is to blame the ideas. The ideas are almost never the problem. Adding people to the creative team rarely moves the number for long, which is a strange result if ideas were the constraint.

Watch where a week actually goes and the bottleneck is production, not imagination. A concept has to be researched against what competitors are running. It has to be written into a script. It has to be shot or generated, then voiced, then assembled, then reviewed against brand rules, then trafficked into the ad account. Each of those steps is a handoff, and each handoff is where the work sits and waits: for the editor with capacity, for the reviewer with time, for the file to move from one tool to the next. Ten a week is not the limit of anyone's creativity. It is the throughput of a pipeline that is mostly made of waiting, and you cannot brief your way past it. This is why hiring another creative rarely helps for long. A second pair of hands adds capacity to one station on the line, but the line still has the same handoffs, the same waiting, and the same single reviewer at the end. You feel faster for a month and then settle at a new ceiling that looks a great deal like the old one.

The move that breaks the wall is to treat creative production as an operable pipeline rather than a craft performed end to end by one person. We built exactly this for our own ad platform, and the shape of it is thirteen named states: a brief moves from competitor analysis to script, to generation, to voiceover, to assembly, to review, to approval, to launch, to measurement, with every transition defined and owned. Naming the states is what makes them parallelisable. When 'make the ad' is one undifferentiated act performed by a person, you get one ad at a time at the speed of that person. When it is thirteen states with clear boundaries, work can be in many states at once, and the machine can own the states that never needed a human. Naming them does one more thing: it tells you exactly where any given brief is stuck at any moment, which is impossible when the whole job is one opaque act in someone's head. A pipeline you can see is a pipeline you can unblock, and a surprising amount of the speed comes simply from never letting work sit invisibly between two steps.

Generation is where model plurality matters most, and it is the part people most often get wrong by betting on a single model. Our pipeline routes across eleven different video generation models, not out of a taste for complexity, but because each has a personality: one is right for product shots, another for motion, another for people, and any one of them can have an off day on a given brief. Routing between them is a decision the system makes per brief. That is what turns 'the model produced something unusable today' from a blocked week into a routing event that costs nobody anything. A pipeline that depends on one model inherits that model's bad days. A pipeline that routes across eleven does not. It is the same instinct that puts a human on the consequential decision, applied to the generative one: no single component, model or person, should be a point of failure the whole week depends on. Redundancy at the noisy step is part of what keeps the reliable steps reliable.

None of this would be safe without the gate, and the gate is what makes the speed usable rather than dangerous. Nothing launches to a live ad account until a human has signed off on it. Five times the output of unreviewed creative is not a growth story; it is a brand-risk multiplier running at speed. So the approval sits exactly before launch, in the tools the team already works in, and the automation handles the twelve states around that one human decision rather than the decision itself. The person is not slowing the pipeline down. The person is doing the one job in the pipeline that has to stay human, while everything upstream and downstream of it runs at machine speed.

The result of building it this way was a fivefold lift in weekly ad output, and it is worth being precise about where that came from. It did not come from better prompts or a cleverer model. It came from naming thirteen states so the work could move in parallel, from routing across eleven models so no single one could stall a week, and from placing exactly one human gate where the brand risk actually lived. The generative model, the part everyone wants to talk about, was the loudest component and close to the least decisive. Volume turned out to be a property of the system, not of any single tool inside it. That is the part worth carrying past advertising. Wherever output is capped and everyone assumes the cap is talent, look first at the handoffs. The ceiling that feels creative is usually operational, and operations is something you can actually fix.

If your ad pipeline is stuck around ten a week, the fix is almost certainly not a better creative team or a newer model. It is to stop treating creative as one heroic act and start treating it as a pipeline: name the states, let the machine own the ones that never needed you, route around any single model's weak days, and put the human sign-off precisely where a mistake would reach a customer. Do that, and the ceiling that felt like a limit on imagination turns out to have been a limit on plumbing all along.

### Private AI, in-house: the shift to internal models and agentic infrastructure

- **URL:** `https://phavella.com/thinking/private-ai-in-house`
- **Published:** 2026-07-11
- **Dek:** A growing set of regulated institutions are moving AI behind their own walls: open models, agentic infrastructure, and a human on the consequential call. The constraint was never the model.
- **Tags:** private AI, self-hosted, agents, agentic infrastructure, data residency, finance, healthcare, government

Something is shifting in how the most regulated organisations buy AI, and it is worth naming plainly. A growing set of large institutions in finance, healthcare, and government are pulling AI in-house: running open models on their own infrastructure, behind their own walls, wired into agentic systems they own rather than rent. This is not a security team being difficult. It is a considered response to a constraint that was always there and is only now satisfiable. A bank's transaction records, a hospital's patient files, a government agency's case notes: for these, sending the data to a vendor's endpoint over the internet is not a preference to be weighed against convenience. It is a line the regulation, the contract, or the board will not let anyone cross. The data cannot leave the building, and until recently that meant the good AI could not come in. That is the part that changed.

What changed is that open models became good enough for the operational work, which is most of the work. Not necessarily for the hardest frontier reasoning, but for the tasks that actually fill an institution's day: classify this document, extract these fields from that form, retrieve the right passage out of millions, draft a reply a person will check before it sends. The interesting question was never the benchmark score of the largest hosted model. It was whether a competent model, wrapped in a serious system, holds up in production. And that question we can answer from experience, because the systems that answer it were never really about the model in the first place. This is the quiet reframing under the whole shift. For years the assumption was that the frontier model was the product and everything around it was plumbing, so of course the good AI seemed to live wherever the best model was hosted. Invert that, treat the model as one replaceable part inside a system you own, and the geography of the whole thing changes.

Take a regulated decision system at scale. One of our principals ran fraud decisioning at a bank that makes sixteen million decisions a month, and has for more than five years. What kept it reliable was never a single model call: it was the thresholds, the review lane for the uncertain cases, the audit trail, and the evaluation that caught drift before customers felt it. Every one of those is a property of the system, not the model, and every one of them runs perfectly well on infrastructure the institution owns. A system that has to hold for five years and stay explainable to a regulator is, if anything, a more natural fit for private infrastructure than for a vendor endpoint that can change under it between one week and the next. The properties a regulator cares about are precisely the ones you would rather not rent. Explainability, stability, and a decision you can reconstruct are easier to guarantee on hardware you control than on a service that can be updated, deprecated, or repriced on someone else's schedule.

Retrieval tells the same story from a different angle. One of our principals built a thirty-million-document search index at a national newsroom, with production applications sitting on top of it. An index of that shape is infrastructure, not magic: it runs inside a private network exactly as well as anywhere else, because retrieval quality was always an evaluation and extraction problem, never a question of which hosted API you called. Once you accept that the hard parts of these systems (the evaluation set, the retrieval layer, the thresholds, the human gate) live outside the model, the case for keeping the whole thing inside your own walls stops being a compromise and starts being the obvious architecture. It stops feeling like giving something up. An institution that already runs its own data centres, its own identity, its own audit and backup is not taking on a strange new burden by hosting the model too. It is extending a discipline it already has to one more component, and getting a system it can inspect end to end in return.

The word doing new work in these conversations is 'agentic,' and it deserves a concrete meaning rather than a hopeful one. An agentic system is a model given tools and a loop: it takes a request, retrieves what it needs, calls the model, invokes a tool to actually do something, and repeats until the task is done. Pulled in-house, every step of that loop runs on the institution's own hardware. The agent retrieves from the institution's own index, calls a model hosted in the institution's own environment, and invokes tools that touch the institution's own systems. The consequential step, the one that moves money or changes a record or sends something to a patient, routes to a human gate before it executes. And the audit trail of the whole loop lands on the institution's own storage, not a vendor's. That is what agentic infrastructure behind your own walls actually means, step by step.

We will not pretend private AI is free, because the honest version is what makes it trustworthy. Running your own models means you carry the inference infrastructure, the evaluation harness, and the on-call that a hosted API quietly carried for you. For a team sending a few thousand requests a month, on data that could live anywhere, a hosted key is the right answer and we will say so. Private AI earns its cost in three situations, and mostly only these three: when the data genuinely cannot leave, when the volume makes per-token pricing punitive at scale, and when the model has to be a stable, owned artifact rather than a dependency that shifts underneath a system that must stay explainable. Outside those cases it is over-engineering. Inside them it is the only thing that actually works, and the institutions now building it know exactly which case they are in. The clarity matters more than the conclusion. A team that has honestly worked those three questions and landed on a hosted key has made a good decision; a team that went private to look serious, without the constraint to justify it, has bought cost it will come to resent. The point is to match the architecture to the real constraint, not to the mood of the moment.

None of this is a new philosophy. It is the same thesis we apply everywhere, moved behind the client's own walls: layer AI onto the systems the organisation already runs rather than replacing them, keep a human on the decisions that are expensive to get wrong, and evaluate before anything reaches production. The only thing that moves is where it all lives. The model that makes the call, the human who signs off, the record of what was decided: on the institution's hardware, inside its network, under its control. For finance, healthcare, and government, that is not a diminished version of AI they are settling for. It is the version they can actually deploy, and the reason the demand is moving in-house is simply that, at last, it can.

### Document intake in an accounting practice: what actually eats the hours

- **URL:** `https://phavella.com/thinking/document-intake-for-accounting-practices`
- **Published:** 2026-07-28
- **Dek:** Client records arrive in every format a client feels like sending, and somebody re-keys them. A system can read the pile and file it. A qualified person still signs anything a regulator will read.
- **Tags:** accounting, document intelligence, intake, sign-off

Ask a partner in a small accounting, audit or tax practice where the year goes and the answer is rarely the accounting. It is the pile in front of the accounting. One client emails a PDF bank statement. Another photographs a stack of receipts on a kitchen table. A third sends a spreadsheet laid out the way a bookkeeper liked it years ago, with a merged header row and the tax column sometimes blank. A fourth posts an envelope. None of them are being difficult. They are sending what they have, in the format they have it in, and all of it lands on the same desk, where somebody opens it, reads it, and types it into the practice's system.

That typing is the real cost, and it appears on no invoice. It shows up instead as the reason a short job takes all afternoon, as the reason the deadline weeks are a wall, and as the reason the person hired to exercise judgement spends the first stretch of every morning doing transcription. It has a second cost that arrives later. Every re-keying is a fresh chance to transpose a digit, and a transposed digit in a filing is not an inconvenience. It is a correction, an apology, and sometimes a cost the practice quietly absorbs to keep a client. The intake pile is where both of those costs are made.

Reading a document and putting its fields somewhere is the thing current models are genuinely good at, and it is worth being precise about why, because the reason is unglamorous. The task is narrow. The model is not being asked to know tax law. It is being asked to look at a page, find the date, the amount, the counterparty and the treatment the client has already indicated, and put each one into a field with a pointer back to the place on the page it came from. That is a bounded task with a checkable answer, which is the shape of work that automates cleanly. Anything broader than that is a different claim, and a much weaker one.

Routing is the other half, and in a practice it matters as much as the reading. A document that has been read still has to arrive somewhere: attached to the right client, in the right period, in the right ledger, flagged if it duplicates the one that client already sent last week. Most of the friction is not any single document. It is the pile's disorder. A system that reads each item, files it against a client and a period, and states plainly what is still missing before a deadline turns a pile into a queue. A queue can be worked through. A pile can only be dreaded.

What both halves have in common is that being wrong is cheap and caught early. If the model reads a nine as a four, the person reviewing sees it beside the source image and fixes it in a second. If it files a receipt into the wrong period, the review catches it long before anything is submitted. None of these failures compound, and that is the real test of what is safe to let a machine do routinely: not whether it is clever, but whether a mistake is visible and reversible at the moment it is made.

Now the line, which is the more important half of this. Extraction is not an opinion. A filing is. The moment the work stops being what does this document say and becomes how should this be treated, it has left the machine's territory. Whether a cost is capital or revenue, whether a transaction sits inside a scheme, what a set of accounts asserts, what an auditor is willing to put a name to: these are professional judgements with a qualification behind them. A model can assemble the evidence and state the question. It cannot carry the responsibility, because responsibility is not a capability. It is an accountability, and it belongs to a person with a licence and a duty.

So the boundary is built in rather than mentioned in a footer. The system reads and drafts; a qualified person reviews and signs. Every number that will end up in a return or a set of accounts passes a human who can see the source beside the entry. Every approval is recorded, so months later it is possible to say who signed what, and on what evidence. This is not a hedge against a weak model. It is the design. In work a regulator, a client or a court may eventually read, the sign-off is the product, and the automation exists to make that sign-off fast rather than to remove it.

We should be straight about where our evidence comes from. We have not run this inside an accounting practice. What we have run is the same shape of work at volumes a practice will never see, which is the more useful thing to know, because the question worth asking is whether the mechanism holds when the pile gets large. One of our principals spent five years building document intelligence at a national newsroom: a search index of thirty million documents, fifteen production applications sitting on top of it, and, measured against the prior workflow, recovery on the order of five hundred hours in every thirty-day window (the exact figure was around four hundred and ninety-seven).

Two other pieces of that record speak more directly to the intake problem. At a private-markets data platform the input was more than a hundred and fifty fund statements from more than six venture and private-equity firms, every one laid out differently, and the job was to read hierarchical tables reliably enough that the numbers could be trusted downstream instead of re-keyed by an analyst. That is the mixed-format problem exactly: the same kinds of fact, arriving in a hundred and fifty different shapes. At a clinical-intelligence company the work was standardising more than a hundred thousand clinical codes so a document's contents mapped onto a controlled vocabulary the profession recognises. A practice has its own controlled vocabulary too. The chart of accounts, the tax codes, a client's own naming for their own costs: mapping a messy document onto a fixed scheme is the same engineering in a different scheme.

None of that makes us your accountants, and we will not pretend otherwise. It means the part of the problem that is engineering has been solved well above a practice's scale, and the part that is professional judgement was never ours to take. What will be new is your last mile: your clients' habits, your filing calendar, the exceptions your seniors carry in their heads and have never written down. That part does not generalise from anyone else's build, and we would rather walk it with you than assume it.

What a practice actually gets is quieter than the pitch usually sounds. The inbox stops being a pile of attachments somebody opens one at a time and becomes a queue of drafted entries with the source on screen beside each one. Chasing becomes a list the system maintains: which clients still owe what, with the reminders drafted and a person sending them. The junior stops transcribing and starts reviewing, which is faster and much closer to the work they came to learn. And the partner still signs, on every number that leaves the building, with the evidence in front of them and the record of that signature kept.

That is the honest version. Nobody is getting an autonomous practice, and anyone offering one has not thought carefully about what happens the first time it is confidently wrong about a filing. What we would take off the desk is the reading, the filing and the chasing. What we would leave exactly where it is, is every judgement that carries a qualification, now reached faster because the paperwork arrives already assembled and already checked.

## Start

Tell us the use case, big or small, custom or straightforward. A principal reads every brief. `https://phavella.com/start`

## Company

- **Website** · https://phavella.com
- **Contact** · https://phavella.com/start
