Call Analyzer
Desktop app · AI call review · Alexey Sukhariev · Design engineer
A sales team makes hundreds of calls a week and the manager hears a handful. Call Analyzer reviews all of them on the manager’s own machine — scoring every call and writing the coaching note, without sending a client’s voice to anyone’s cloud. I built it end to end: product, design system, backend, prompts, installer. A batch now finishes without anyone watching it, with failures under 1%.
The calls nobody hears
A sales team makes hundreds of calls a week. Their manager listens to a handful, usually after a number on a dashboard has already gone wrong — and two managers reviewing the same call rarely score it the same way. Call Analyzer takes the whole pile: it transcribes each call, scores it against the criteria that team actually cares about, and writes two things a human can act on — a message for the agent and a note for the manager. The decision that shaped everything else: it runs on the manager’s own computer. Not a SaaS, not a web account — a Windows application with a local database and a local model, because the calls carry client data that cannot be handed to someone else’s cloud.
Three constraints wrote the architecture
Buying something off the shelf failed on three counts, and each one became a rule. Confidentiality: recordings cannot leave the machine, so the database and the default model are local and the cloud is an opt-in for the cases a small model cannot handle. Cost at volume: hundreds of calls a week through a paid API is a bill that grows with success, so the local model has to be good enough to be the default. And no IT department: the user is a head of sales, not an engineer, so anything harder than download-and-double-click would never be used. That last sentence is what turned a working web app into an installer, an onboarding flow and a hardware check.
What the reviewer actually gets
One call, one screen. The score and the verdict sit at the top, the criteria breakdown explains where they came from, and the goal line says what this conversation was supposed to achieve. Nothing here is a black box: every number traces to a criterion, and every criterion traces to something that was said.
Evidence, then action
Underneath sits the evidence — strengths, top errors, missed opportunities, and the lines worth keeping — each tied to a timecode, so any claim can be checked in seconds. Then the two outputs that make the tool useful rather than interesting: a message ready to send to the agent, written as coaching rather than judgment, and a separate note for the manager. Same analysis, two readers, different jobs.
A page for every agent
A single call is an anecdote; the agent record is where it becomes a pattern. Score over time, the errors that keep repeating with their frequency, the strengths that show up again and again. A manager preparing a one-to-one opens this page and already knows what the conversation is about.
A prompt is a contract with a model class
The most expensive mistake of the project was an assumption, not a bug. After an import every agent’s scoring came back empty: analyses present, status OK, overall score null across the board. The prompts had been iterated against a 32B-class model in the cloud, with reference reviews as the gold standard — but an 8 GB card fits an 8B model, four times smaller. It filled two criteria out of five and gave up; the score function required a full set, returned null, and the emptiness propagated to every screen. So the prompt stopped being a property of the review scenario and became a property of the pair — scenario × active model — with a canonical key so the same model under two providers resolves to one prompt. The arithmetic moved out of the prompt into code, because a small model cannot be trusted to compute a weighted average. And the iteration method became a rule of its own: one surgical change per round, measured against the reference set, because stacking changes made the small model collapse every score to a flat 5 out of 10. On the local 8B model, first calls now reach 12/12 on call status, 11/12 on verdict and 9/9 on completeness, with a mean error of 0.44 on the overall score. Retention scores worse — 5/8 on verdict — and that is written down rather than hidden.
A model does not keep its word
Some analyses came back with an empty radar. The model had moved the scores to the top level of the JSON instead of the nested object, and sometimes renamed a key to match the wording of the instruction — objectionHandling where the schema said objection. The fix is a tolerant parser: canonical key first, then the top level, then a synonym map per criterion; and on top of it a post-processing layer for the judgements the model got wrong systematically, such as rewarding a polite refusal that resolved nothing. The rule that came out of it: a JSON schema inside a prompt is a wish, and the contract lives in the parser.
Four layers against a rate limit
Re-analysing 199 calls through a free cloud tier produced 65 failures — a third of the batch. The limit was 6,000 tokens a minute and one review costs 5,000–7,000; pausing sixty seconds between calls did nothing, because the window is rolling, so requests made at the end of one minute are still inside it at the start of the next. What worked was defence in depth: a throttle between calls, a 90-second pause and retry on a limit error, 120 seconds on the second, and a background sweep every five minutes that returns rate-limited reviews to the queue — with a cut-off so it cannot pick up the errors it just wrote. Failures dropped below 1%, and a batch now finishes without anyone watching it.
A queue that survives the process
Three bugs with one root cause: the queue lived in memory, and processes die. A restart lost the queue while the database still said running, so calls hung forever — fixed by recovery on startup. The first version of that recovery re-queued every pending call for each unfinished job: four stale jobs × 199 calls = 796 entries for 199 real calls, and the counter froze at 0 of 199 while the machine worked at full tilt — fixed by deduplication, newest job wins. And one call that could not fit the token limit failed, was retried, failed again, forever, with the pending count oscillating between 21 and 22: a queue that looked alive and moved nowhere.
The scoring scenarios belong to the user
The first version had two hard-coded review types, which is fine until the team wants a third. So scenarios became data: a pipeline is a name, a colour, a set of weighted criteria and a prompt, created in the interface and duplicated from templates. Calls arrive by CSV or by watched folder — drop an audio file in and it is transcribed locally and analysed under that pipeline’s criteria without anyone opening the app.
From a web app to something you can install
The stack includes a model engine and a model file of 2–20 GB; neither can be baked into an executable, and the user has no IT department. So the app installs its own engine, downloads the model with a live progress readout, and — the part I like most — explains the user’s computer to them: it reads the machine, finds the graphics card, and recommends the model that will actually fit in video memory. That logic was wrong at first: it recommended by RAM, when for GPU inference the bottleneck is video memory — if the model does not fit, layers spill into RAM and inference slows by an order of magnitude. Shipping was a set of deliberately unglamorous decisions: a native window wrapping the local server instead of rewriting fifteen screens in a native toolkit; read-only resources separated from user data, so reinstalling never costs the call history; a per-user installer that needs no administrator rights. Licensing runs without a server of my own — a signed payload with a device fingerprint, verified against a key built into the client, with a signed offline cache for the days the licence server is unreachable.
Three design systems in seven weeks
The first interface was cyberpunk — neon, scanlines, cut corners, uppercase headers. It photographed well and became tiring in the second minute of real use, and was thrown out with a formulation I kept: a hint, not a costume. The second was Apple HIG by the book. The third, the one that shipped, is built on tokens pulled from the official iOS 26 Figma kit — real colour, material and type values exported into CSS variables instead of numbers copied from blog posts. Eleven type classes, a 4-point spacing grid, four levels of button prominence, translucency reserved for floating surfaces only. The insight that made the difference: the onboarding screen looked foreign not because it was badly built, but because it still used hard-coded colours while every other screen had moved to tokens. Two things I measure rather than claim. Accessibility went from 8/100 to 80/100 — tertiary text contrast from 2.93:1 to 5.7:1, decorative icons hidden from screen readers, accessible names on icon buttons, a skip link, reduced-motion support. And performance is part of the design: the settings page took 840 ms to open because it checked three providers in sequence; parallel checks plus a short-lived cache brought it to 7 ms warm and took 280 ms × 12 times a minute off the background polling. Progress used to update by reloading the page every 3.5 seconds, which threw away scroll position, focus and anything typed into a form; now only the nodes whose values changed are touched.
The app as a tool for another AI
Late in the project the application started exposing its own MCP server: twelve tools over JSON-RPC, so an external AI client can read calls and agents, trigger an analysis and write its own verdict back through the same path the internal worker uses — external judgements stay consistent with internal ones. The transport is hand-written rather than an SDK, deliberately: one fewer dependency in a fragile single-file build. It has a useful side effect, too — a strong external model reaches exactly the calls the small local one grades poorly.
Where it stands
A working product: built executable, per-user installer, server-side licence activation, MCP integration, tested end to end on a separate machine. It has not been released publicly, and the honest edges are worth stating. Retention on the local 8B model is weaker than first calls — it does not recognise failure by inaction, the polite conversation where nothing happens. There are no automated tests; everything was verified by hand, which is the real limit on how fast new features can land. Part of the accessibility list is still open. The single-file build starts slowly because 114 MB unpacks on every launch — a trade accepted for one file to send. Self-updating is designed and not built. And the licence scheme stops casual copying, not a patched binary: it closes exactly the case it was built for and nothing more.
← Alexey Sukhariev