A Real-Time AI Meeting Assistant for the Enterprise, and the Road to an AI Phone Receptionist

Most "AI meeting notes" products work after the fact: the call ends, a bot that joined the meeting sends a summary an hour later, and the useful moment has already passed. I wanted the opposite: an assistant that listens during the call, knows the company's products and the customer in front of you, and puts a usable answer on screen within a couple of seconds of a question being asked. This post covers what I built, the stack behind it, how it fits a sales team today, and why the same engine is a short step away from an AI phone agent that can answer the front desk and take bookings.

The problem I wanted to solve

Watch a sales rep on a discovery call and you see the same pattern again and again. The customer asks something specific: "Does this integrate with our ERP?", "What's the price for 40 seats on annual billing?", "How is this different from what we have now?" The rep either knows, or says "let me get back to you", or puts the customer on hold to search the product wiki. Afterwards someone has to type notes into the CRM, and half of the detail is lost.

Three things would fix most of that:

  • A live transcript that knows who said what, without a bot joining the meeting.
  • An answer on demand, grounded in company material (product sheets, price rules, the account's context), written so the rep can say it out loud.
  • Clean output at the end: a transcript and summary that can go into the systems the business already runs.

What it does today

The result is a Windows desktop app. It works with any meeting tool (Teams, Zoom, Google Meet, Slack huddles, a softphone) because it listens to the computer's audio rather than integrating with one platform. Nobody has to invite a bot, and IT does not have to approve a new meeting add-on.

  • Two-channel live transcript. Speaker audio is labelled Others and the microphone is labelled Me. Each channel has its own recognizer, so the speaker label comes from the audio source rather than from a guess. Lines from both channels are ordered by when people started speaking, so the conversation reads correctly even when one channel finishes recognizing later than the other.
  • Draft answer. One hotkey finds the latest question aimed at you in the transcript and drafts a spoken-style answer. You can also click any line to answer that exact turn, or type a question that was misheard.
  • Draft + screen. The same, plus a capture of what is on screen: the slide the customer is sharing, a CRM record, a price sheet. The app also reads the window's text through Windows UI Automation, so numbers, names and code reach the model exactly as written instead of being guessed from pixels.
  • Chime in. When nobody is asking you anything, the assistant suggests three ways into the discussion: something to add, a new angle, and a risk or objection to raise.
  • Auto mode. While listening, the app spots a new question from the other side (a question mark, or an opener like what / how / can you), waits for the speaker to finish, and drafts the answer without a key press.
  • Answer styles. Bullets, natural paragraph, one-liner, a short story-style answer, or "ask back", which suggests two or three smart questions to return to the customer. That last one is the most useful on discovery calls.
  • Company context. A context panel holds standing instructions (up to 30,000 characters) such as "We sell X; enterprise tier includes SSO; never quote discounts above 15%", plus a per-meeting note like "Renewal call with Contoso, champion is the IT manager". Presets switch between context sets in one click.
  • Phone mode. For calls taken on the PC (Phone Link, Teams Phone, softphones) or on a speakerphone next to the mic. In speakerphone mode both voices arrive on one microphone, so the app uses Azure Conversation Transcription to separate speakers, and you can relabel a voice with a right-click if it guesses wrong.
  • Mini bar. A small always-on-top strip with listen, Auto, draft and the latest answer, so the assistant does not cover the meeting window.
  • Output. Copy or save the full transcript (up to 3,000 lines kept in the session), ready to summarize or paste into the CRM.

The stack

LayerChoiceWhy
App.NET 9, WPFNative Windows: system-wide audio, global hotkeys, tray, per-monitor DPI.
Audio captureNAudio, WASAPI loopback + microphoneHears every meeting app without integrating with any of them.
Speech-to-textAzure AI Speech (continuous recognition and Conversation Transcription with diarization); Windows System.Speech as an offline fallbackCloud accuracy with speaker separation, plus a free offline option.
LLMClaude via the official Anthropic C# SDK, streamingStrong reasoning over long context and images; tokens render as they arrive.
Screen understandingScreen capture + Windows UI Automation textThe model gets the exact text of the CRM or document on screen, not just a picture.
RenderingMarkdig.WpfAnswers show as Markdown: tables, lists, code.
SecurityWindows DPAPIAPI keys are encrypted per Windows user, never stored as plain text.

The data flow is short on purpose:

 Meeting app audio ──► WASAPI loopback ─┐
                                        ├─► 16 kHz mono PCM ─► Azure Speech ─► Live transcript (Others / Me)
 Microphone ─────────► WASAPI capture ──┘                                         │
                                                                                  ▼
 Screen + UI Automation text ─────────────────────────────►  Prompt builder (context, style, language)
                                                                                  │
                                                                                  ▼
                                                              Claude (streaming) ─► Answer card / Mini bar

Engineering notes that mattered

Latency is the product

An answer that arrives eight seconds after the question is a meeting summary, not an assistant. Several small decisions add up:

  • Streaming everywhere. The first words show up while the model is still writing. The view re-renders at most about twelve times a second, so streaming does not choke the UI.
  • A separate, faster model for live calls. Live meetings can use a lighter model with low reasoning effort, while document-heavy work keeps the larger model.
  • Send only what matters. Each draft sends roughly the last 6,000 characters of the transcript, not the whole call. The full transcript stays in the app.
  • Prompt caching where it pays. Long multi-turn sessions cache the system prompt and earlier turns, so later turns read them at a fraction of the price. One-shot drafts skip caching on purpose, because the cache write costs more and would never be read back.

Audio is messier than it looks

  • Meeting apps output float32, 16/24/32-bit PCM, stereo or more. Everything is mixed to mono and resampled to 16 kHz / 16-bit before it reaches the recognizer.
  • WASAPI loopback sends nothing when the PC is silent, so the recognizer never learns that a sentence ended. The fix is to inject short blocks of silence so phrases close on time.
  • Auto mode waits for the speaker to finish (about 1.8 seconds of quiet, at most 5 seconds) and keeps at least 6 seconds between automatic answers. Without this it answers half-asked questions.
  • Stopping a recognizer that has hung must never freeze the app, so shutdown runs on the thread pool with a timeout.

Reading the screen like a person, safely

UI Automation gives the text of documents, browser pages, editors and CRM screens, including parts scrolled out of view. Reads are capped by time and size so a busy app cannot stall the assistant. Password fields and input boxes are never read, and the user sees the exact text that will be sent, with a one-click "don't send". Windows on a blocklist (banking, HR systems, anything sensitive) are never captured.

Prompts as data, not code

All prompts live in one JSON file embedded in the app. The code only picks variants (language, answer style, transcript type) and fills placeholders in a single pass, so text pasted into the context panel cannot inject new placeholders. Because answers follow a fixed format (detected question, topic, or "no question"), the UI can render them as structured cards instead of a wall of text.

A sales showcase

Here is how this plays out on a typical B2B sales call.

  1. Before the call, the rep picks the "Sales" context preset: product sheet, pricing rules, competitor notes, discount limits. They add one line about the account.
  2. During discovery, Auto mode is on. The customer asks, "Can we keep our existing SSO?" Two seconds later the mini bar shows three bullets: yes on the Enterprise tier, the supported providers, and a follow-up question about their identity provider.
  3. On pricing, the customer shares a spreadsheet of their current costs. The rep presses Draft + screen. The answer reads the numbers off the shared sheet and frames a comparison within the price rules from the context panel.
  4. On objections, "We already have a tool for this", the rep presses Chime in and gets a differentiator, a question that uncovers the pain point, and a risk worth naming honestly.
  5. After the call, the transcript is saved and summarized into next steps, ready for the CRM.

The rep stays in charge: the assistant suggests, the human speaks. New reps sound like they have been on the team for a year, and experienced reps stop losing the details.

Connecting to the rest of the business

Today the integration points are deliberately simple: the assistant reads whatever system is on screen (CRM, ERP, ticketing, a shared spreadsheet) through capture and UI Automation, and the transcript leaves the app as text. That already works with every system without an API key or an IT project.

The next layer replaces "read the screen" with direct tools. Claude supports tool use, and the same idea is packaged as MCP servers, so the assistant can call business systems instead of only looking at them:

  • CRM (HubSpot, Salesforce, Dynamics): look up the account and open deals at the start of a call; after the call, write the summary, next steps and updated deal stage.
  • Product and pricing: quote from the real price book instead of a pasted sheet.
  • Calendar: propose and book the follow-up meeting while everyone is still on the call.
  • Ticketing and tasks (Jira, Asana, Linear): turn action items into tickets with owners.
  • Knowledge base: answer from the company wiki and past proposals, with sources.

The rule I would keep from day one is that reads can be automatic, but every write to a business system shows a preview and waits for the user's confirmation.

Where this goes next: an AI phone agent

Strip the desktop UI away and what is left is a pipeline: audio in → transcript with speakers → reasoning with business context → response. That is also the core of an AI phone agent. The difference is that the response is spoken back to the caller and the actions are real.

 Caller ─► Telephony (SIP / Twilio / Azure Communication Services)
              │ audio stream
              ▼
          Streaming speech-to-text ─► Claude + tools ─► Text-to-speech ─► Caller
                                         │
                                         ├─ check availability (booking system / POS)
                                         ├─ create or change a booking / order
                                         ├─ look up the customer (CRM)
                                         └─ hand off to a human with a summary

Use cases I am designing for:

  • Front desk / receptionist for clinics, salons, spas and offices: answers opening hours, services and prices, and routes calls, 24/7 and in more than one language (English and Vietnamese are already supported in the transcription layer).
  • Bookings and reservations for restaurants, nail salons, repair shops: checks open slots, books, reschedules and sends an SMS confirmation.
  • Phone orders: takes an order, confirms the items back to the caller, and pushes it into the POS.
  • After-hours and overflow: answers when every staff member is busy, then hands the call over with a written summary so the customer never repeats themselves.

What carries over directly from the meeting assistant: the audio handling, the end-of-turn timing (when has the caller actually finished?), speaker separation, the latency budget, the prompt layer, and the guardrails around what gets sent where. What is new: barge-in (the caller interrupts the agent), text-to-speech with a natural voice, telephony, and strict confirmation before any booking or payment is committed.

Privacy and compliance

Any product that listens to conversations has to earn trust first:

  • Inform participants and follow local call-recording laws. For a phone agent, the greeting says the caller is speaking with an AI assistant.
  • Transcripts and screen captures go to the configured AI and speech providers only. Enterprise deployments can use their own Azure Speech resource and Claude API account, keeping data under the company's agreements.
  • Secrets are encrypted per user; sensitive windows can be blocklisted; the user sees what is sent before it is sent.
  • Humans approve every write to a business system.

Takeaways

  • Live beats after-the-fact. The value of an AI meeting assistant is in the two seconds after a question, not in the summary an hour later.
  • Listen at the OS level. Capturing system audio made the assistant work with every meeting and phone tool, with no bot and no add-on to approve.
  • Grounding is the feature. A generic model gives generic answers. Company context, the screen in front of the rep, and later direct CRM access are what make answers worth saying out loud.
  • Build the pipeline once. The same audio → transcript → reasoning → action engine powers a sales copilot today and a phone receptionist tomorrow.

If you run a sales team, a front desk, or a business that lives on the phone and want to try this, get in touch. I am happy to set up a pilot with your own products, price list and booking system.


More case studies on the Featured Projects page, or reach me at letanphp@gmail.com.

Comments

Popular posts from this blog

Featured Projects: Automation and AI Systems I Built and Run

AWS API gateway, S3

Business case: Monitor mailbox and auto-save the attachment to a SharePoint