<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
    <channel>
        <title>Moss Blog</title>
        <link>https://mossdev.work/blog</link>
        <description>Insights, tutorials, and updates from the Moss team on real-time semantic search, AI agents, voice AI, and retrieval-augmented generation.</description>
        <lastBuildDate>Tue, 06 Oct 2026 13:58:36 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <language>en</language>
        <image>
            <title>Moss Blog</title>
            <url>https://mossdev.work/Favicon.svg</url>
            <link>https://mossdev.work/blog</link>
        </image>
        <copyright>All rights reserved 2026, InferEdge Inc.</copyright>
        <item>
            <title><![CDATA[Why We Built a Search Runtime in Rust and Compiled It to WebAssembly]]></title>
            <link>https://mossdev.work/blog/rust-webassembly-search-runtime</link>
            <guid isPermaLink="false">https://mossdev.work/blog/rust-webassembly-search-runtime</guid>
            <pubDate>Fri, 25 Sep 2026 09:00:00 GMT</pubDate>
            <description><![CDATA[Why Moss runs retrieval inside the agent's runtime: how index building is split from queries, why Rust and WebAssembly fit an embedded search engine, what happens when an index loads and changes, what the published benchmarks show, and where portability adds engineering work.]]></description>
            <content:encoded><![CDATA[
We built Moss around the idea that retrieval belongs in the runtime with the agent. When an agent needs knowledge to answer a question, it should be able to search that knowledge where it is already running, whether that is a browser tab, a server, a phone, or an edge environment. We wanted retrieval to be a capability the agent carries with it, without adding a network service to every lookup.

That decision came from the applications we were building for. A voice agent checking a refund policy has to retrieve the answer while a caller waits. A website assistant may need to search inside the browser, and a field application may need access to its knowledge after the connection drops. These applications have different deployment constraints, but they all benefit when retrieval can live beside the code that uses it.

Rust and WebAssembly followed from that requirement. Rust gave us control over execution and memory in an embedded library, while WebAssembly made that library available inside the browser. The more consequential design choice was deciding which work belonged in the cloud and which work had to happen in the agent’s process.

## How Moss Separates Index Building from Queries

Moss Cloud handles the work of preparing and distributing an index. It ingests documents, builds the searchable knowledge, and makes updates available to the runtime. The application loads the index locally before it starts querying. Once it is loaded, the runtime can embed a query, retrieve relevant documents, apply filtering and ranking, and return context to the agent without a retrieval request to the cloud. The offline search documentation describes this lifecycle.

This split keeps index maintenance available as a managed service while making query execution part of the application. It also changes the failure boundary. After the required assets and index are loaded, a temporary loss of connectivity does not require every knowledge lookup to fail, although receiving new content still requires synchronization.

The diagram shows both parts of the architecture. The build targets determine how the Rust core runs on each platform. Index delivery is a separate path from Moss Cloud to the deployed runtime, where the agent and its loaded knowledge meet.

<figure className="art-figure art-figure--image reveal">
  <img src="/blog/rust-webassembly-search-runtime/architecture.png" width="3200" height="1800" loading="lazy" alt="Two architectures compared. Traditional retrieval: the agent makes a network request to a vector database for each query. Moss architecture: queries run inside the agent's runtime. A shared Rust core is compiled as a WebAssembly build, deployed to the browser, and a native build, deployed to server, mobile, and edge. In each, the agent talks to a local runtime with a loaded index. Moss Cloud builds and syncs indexes, and a dashed path delivers the index and background refreshes to both runtimes." />
</figure>

<p className="art-caption">The solid arrows show compilation and deployment. The dashed path delivers index updates outside the per-query retrieval path.</p>

## Why Rust and WebAssembly Fit the Runtime

An embedded search engine shares a process with the application using it. Its memory usage and execution time therefore become part of that application’s behavior. Rust was a good fit because it lets us manage memory without adding a garbage collector to the search core, while its ownership system prevents many classes of memory errors in safe code. That gives us useful control over latency without making memory safety entirely a matter of programmer discipline. Rust’s ownership model explains the underlying tradeoff.

Keeping retrieval behavior in the Rust core also means a ranking change can reach the supported SDKs through the same implementation. Platform bindings handle how applications call the runtime and receive results. They do not need to recreate the search logic, which makes it easier to reason about correctness as the runtime moves between environments.

WebAssembly gives the browser a way to execute that core within its own security and execution model. Native builds serve hosts that can load a platform library directly. These targets can expose a similar application workflow, but they still have different startup costs and access to system resources.

Sharing the implementation helps keep behavior aligned; it does not imply identical performance on a laptop, a phone, and a browser tab.

## What Happens When an Index Loads and Changes

A local query starts being useful only after the runtime is ready, so startup deserves its own budget. In the browser, first use includes downloading the WebAssembly module and embedding model, as described in the browser API reference. The application also has to fetch and load its index.

Network transfer, runtime initialization, and index loading all contribute to the time before the first query can run. A query latency measurement by itself does not describe that experience.

Loading an index does not require embedding the document collection again. That work has already happened during the cloud build. Caching can also reduce repeated transfer: the native JavaScript SDK can keep an index on disk and reuse it on a later load when the cloud version has not changed. This avoids fetching the same data again, though the application still needs a ready runtime and a loaded index. The storage documentation covers that distinction.

Index size matters in several ways. The downloaded bytes affect transfer time and local storage, while the loaded index consumes memory alongside the embedding model and the application itself. Text, metadata, and temporary query buffers also contribute to the process’s footprint. Document count alone is therefore an incomplete sizing guide: a collection of short support answers and a collection of long manuals can impose different costs even when they contain the same number of records.

Updates have a lifecycle too. With automatic refresh enabled, the runtime checks for a newer index and downloads a ready version in the background. Queries already in progress can finish against the current version while later calls use the replacement. That behavior, documented in data hydration and sync, keeps refresh work from interrupting an active query.

It also means memory planning should leave room for an update while the old version is still serving requests, rather than treating steady-state usage as the maximum.

## What the Published Performance Results Show

Our repository reports an end-to-end query benchmark over 100,000 documents, with 750 measured queries returning the top five results on a MacBook Pro with an M4 Pro and 24 GB of memory. The measurement includes query embedding and search. The published results give the following latency distribution.

<div className="art-table-wrap">
  <table className="art-table">
    <thead><tr><th scope="col">Workload</th><th scope="col">Median</th><th scope="col">P95</th><th scope="col">P99</th></tr></thead>
    <tbody>
      <tr><th scope="row">Embedding and search</th><td>3.1 ms</td><td>4.3 ms</td><td>5.4 ms</td></tr>
    </tbody>
  </table>
</div>

For that setup, even the 99th percentile remained below six milliseconds. The scope matters: this is a query benchmark, and it does not establish download time, index loading time, or browser performance. It shows that the local runtime can fit retrieval into a small part of the response budget on the tested machine.

A separate voice-agent pipeline profile reports the search processing step inside a hosted retrieval flow and a local one. The hosted database processing step had a median of 18 ms; the in-process search step had a median of 1.2 ms. Those figures exclude query embedding and result processing, so they describe a narrower part of the request than the repository benchmark.

The pipeline profile does not specify its hardware, corpus size, or query count, so it is an illustrative result for that workload rather than a reproducible comparison across platforms. Taken with the repository benchmark, it gives us evidence for local execution while leaving native-versus-WebAssembly performance, startup time, and peak memory to be measured on the application’s actual deployment target.

## Where Portability Adds Engineering Work

The browser makes the cost of distribution immediately visible. A larger module or model increases the work before retrieval is ready, and loading an index consumes resources in a tab that also has to render the interface and run the rest of the application. A deployment decision therefore has to account for the first visit as well as repeated queries. An index that works comfortably beside a server agent may need a smaller scope when it is delivered to a browser on a constrained device.

Browser execution also depends on the page that hosts it. For example, shared memory between workers requires a secure context and cross-origin isolation under the browser’s shared-memory rules. That is a constraint on the capabilities a library can assume when embedded in an arbitrary site.

Serving WebAssembly brings its own integration details: streaming instantiation expects the appropriate content type, and a site’s content security policy can restrict compilation, as the WebAssembly documentation explains. A successful build alone does not establish that the surrounding page can run it correctly.

Native distribution moves the work elsewhere. A native library has to match the operating system, processor architecture, and host language that will load it. Linux compatibility is especially easy to miss when development and deployment environments differ. Keeping retrieval in one core reduces duplicated search code, but each supported package still needs to load and behave correctly in its host environment. Compatibility is part of the runtime’s usefulness, alongside the time it takes to answer a query.

Local execution also makes resource ownership explicit. The application pays for the CPU and memory used by retrieval, and the index it loads needs to fit the device’s budget. Where the corpus is too large to distribute to clients, running Moss beside a server-based agent can be a better placement. The architecture gives us a way to choose that location without changing the application into a client of a separate search service on every turn.

## Keeping Retrieval with the Agent

The engineering work follows from the reason we built Moss this way. Retrieval belongs where the agent runs, with the knowledge it needs available when it needs to act.

Rust gives us the foundation for an embedded implementation, and WebAssembly carries it into the browser. Together with cloud index building and synchronization, they let us keep the work of answering a query close to the agent while managing the knowledge it searches over time.

You can [try Moss](https://portal.usemoss.dev) or explore the [documentation](https://docs.moss.dev) to see how the runtime fits into your application.
]]></content:encoded>
            <author>Ashvath Suresh Kumar</author>
            <category>Rust</category>
            <category>WebAssembly</category>
            <category>WASM</category>
            <category>search runtime</category>
            <category>embedded search</category>
            <category>in-process retrieval</category>
            <category>browser search</category>
            <category>on-device AI</category>
            <category>vector search</category>
            <category>AI agents</category>
            <enclosure url="https://mossdev.work/blog/rust-webassembly-search-runtime/og.png" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[How Aside Powers Persistent AI Memory Across 80,000+ Devices]]></title>
            <link>https://mossdev.work/blog/aside-persistent-ai-memory</link>
            <guid isPermaLink="false">https://mossdev.work/blog/aside-persistent-ai-memory</guid>
            <pubDate>Wed, 23 Sep 2026 09:00:00 GMT</pubDate>
            <description><![CDATA[Aside embeds 113M documents and 24.8B tokens a month on device across 80,000+ devices in 150+ countries. How its AI browser uses Moss for private, local memory retrieval with sub-10ms queries.]]></description>
            <content:encoded><![CDATA[
Aside is building an AI browser that remembers what users do, turning browsing history and past tasks into persistent context that agents can use to take action. The product is growing 25% week over week, with its memory infrastructure now spanning more than 80,000 devices across 150+ countries. Every month, Aside processes approximately 113M documents and 24.8B tokens directly on device, giving its agents access to a growing body of private, searchable context.

Making that work at this scale meant building memory infrastructure that could embed and retrieve large amounts of information directly on users’ devices, without sending sensitive data to the cloud. Moss provides that retrieval layer, including the local semantic indexing, embedding infrastructure, and retrieval engine that sits underneath Aside’s memory system.

I spoke with Jun Kim, Co-Founder and CEO of Aside, about why persistent memory is becoming an important part of AI agents, how Aside built its memory architecture, and why local retrieval became a critical part of the product.

<PullQuote>“The kind of product picking, I believe, is not achievable without a good memory layer. That's why we invest a lot in memory.”</PullQuote>

## Aside’s mission: Give AI agents the memory to understand how people work

AI agents are becoming increasingly capable at completing tasks, but they still need context to understand the work behind those tasks. A user might spend an hour working with an agent on a project, only to have to explain the same context the next time they open it. The agent may know how to complete the task, but it does not necessarily know what the user was working on yesterday, who they were talking to, or what happened the last time they asked it to do something.

Aside is tackling that problem through the browser. Because so much of modern work already happens in the browser, it gives Aside access to a rich source of context without requiring users to connect every tool they use. The browser can learn from browsing history and past tasks, building a persistent understanding of the websites, projects, people, conversations, and work that matter to each user.

As Jun explains:

<PullQuote>“Browser is our daily driver, so we use browser every day, and we do 80% of work on the browser. But at the same time, browser is the most richest context source.”</PullQuote>

The goal is an AI browser that can carry context from one task into the next, rather than treating every interaction as a blank slate.

## 113M documents and 24.8B tokens every month: Building memory on device

Aside’s memory system extracts information from individual interactions and turns it into episodic memory. It organizes that information into different types of context, including people, projects, websites, and user-specific information. When information does not fit neatly into an existing category, the system can evolve its own taxonomy, creating what Jun describes as a “self organizing memory.”

The challenge is figuring out what information is useful, organizing it so an agent can understand it, and retrieving the right context when it matters. Aside built its own Markdown semantic chunker to break individual memory entries into retrievable chunks, then uses Moss for the local embedding, semantic indexing, and retrieval. The team was specifically looking for three things from its memory infrastructure: retrieval quality, local execution, and speed.

Local execution was particularly important because Aside’s memory can contain browsing history, work context, and other sensitive information. The team did not want that data sent to a cloud memory service simply so it could be searched.

<PullQuote>“It has to be local. I didn't want to use cloud back memory at the moment because it is sensitive data.”</PullQuote>

With Moss, Aside keeps memory on the user’s device while giving its agents semantic search across a growing amount of context. Today, that means processing around 113M documents and 24.8B tokens each month across more than 80,000 devices in 150+ countries. The volume is growing 25% week over week as more users rely on Aside to remember their browsing history and past work.

## 80,000+ devices across 150+ countries: Scaling local retrieval

Running retrieval locally creates a different set of infrastructure requirements. Aside needs its memory system to work across a distributed fleet of devices while maintaining the performance and reliability expected from an AI agent that is constantly accessing context. Embedding and semantic retrieval also need to happen locally rather than relying on a centralized cloud database.

Aside considered building its own retrieval infrastructure, but the question was less about whether the team could build it and more about where its engineering resources were best spent. Retrieval infrastructure is a specialized area that spans embedding models, semantic indexing, local execution, and fast retrieval across a wide range of devices. Building and maintaining that infrastructure internally would also mean taking on the ongoing work of optimizing it as Aside’s device footprint and memory volume continued to grow.

The team wanted to focus its engineering effort on the parts of Aside that make the product what it is: the AI browser, agent harness, memory extraction and organization system, and password manager. Moss handles the underlying retrieval infrastructure, while Aside controls how memory is created, organized, and ultimately used by its agents.

## \<10ms query latency: Keeping retrieval out of the way

For an agent, memory is only useful if retrieving it does not slow down the experience. As Aside’s memory system grows, retrieval needs to happen quickly enough that accessing historical context feels like a natural part of using the agent rather than a separate step. Moss provides query latency of less than 10 milliseconds, giving Aside a fast retrieval layer for its growing on-device memory system.

For Aside, the goal is for users to stop thinking about whether the agent remembers something or whether they need to explain it again. The relevant context should simply be available when the agent needs it.

## Turning persistent memory into action

The value of persistent memory becomes more apparent when it changes what the user has to tell the agent. During our conversation, Jun showed me a scenario involving an investor who had asked for a financial and investor update. Instead of explaining who the person was, finding the relevant email thread, and describing what needed to happen, Jun simply told Aside to handle what the investor had sent.

Aside remembered who the person was, found the relevant email thread, understood the surrounding context, and began coordinating the work.

<PullQuote>“I was really surprised, because Aside handled the exact same thing I wanted with only four words.”</PullQuote>

That is what makes persistent memory useful in practice. The agent is not just storing more information; it is reducing the amount of context the user has to provide every time they want something done. Because Aside is a browser, that memory can also connect directly to action. Users do not necessarily need to connect Gmail, Google Drive, GitHub, Google Calendar, Slack, and other individual tools before asking the agent to work across them. The browser already has access to the websites and credentials needed to carry out the work.

<PullQuote>“You don't have to connect anything, and you don't have to type a long prompt thanks to memory layer and our password manager credential system.”</PullQuote>

<Screenshot
	src="/blog/aside-persistent-ai-memory/memory-history.png"
	width="2000"
	height="1005"
	alt="Aside’s Memory History screen, under the heading “See what Aside remembers”: memory is written in markdown and stored on device. A list of memory updates extracted from sessions (each with its time, duration, token count and files changed) sits beside the selected update’s diff to episodic/2026-06-21.md and the memory extraction subagent’s prompt."
/>

## Four words instead of a paragraph

Aside does not treat memory as the only source of truth. If a memory retrieval does not return exactly what the agent needs, it can continue searching for context by opening the relevant website and verifying the information directly. Memory gives the agent a starting point: it helps it understand the user’s history, identify where relevant information may exist, and decide what to look for next.

That means memory does not have to contain every piece of information forever. It needs to give the agent enough context to understand what the user means and take the next step without making the user start from scratch. In practice, that moves the interaction away from lengthy prompts and toward simply telling the agent what needs to get done.

## A local memory layer for AI agents

Memory is becoming a core part of how Aside works. The company has built its own browser and agent harness, memory extraction and organization system, and password manager, while using Moss for the infrastructure underneath memory retrieval.

That separation lets Aside focus its engineering effort on the parts of the product that are unique to its approach, while relying on specialized infrastructure for embedding, indexing, and local retrieval across a growing number of devices. Keeping the memory on device also means sensitive user information does not need to be sent to a cloud memory service simply to make it searchable. Jun said this has opened up interest from companies operating in areas with stricter security requirements, including finance and law.

For Aside, memory is becoming an increasingly important part of what makes an AI browser useful.

<PullQuote>“The kind of product picking, I believe, is not achievable without a good memory layer. That's why we invest a lot in memory.”</PullQuote>

Today, Moss provides the infrastructure underneath that memory system, helping Aside embed and retrieve 113M documents and 24.8B tokens every month across more than 80,000 devices while keeping memory local to the user.

## Built with Moss

[**Aside**](https://aside.com/) is building an AI browser that remembers users’ work, turning browsing history and past tasks into persistent context that agents can use across future tasks.

**Moss** provides the local memory retrieval infrastructure behind Aside, including high-performance embedding, semantic indexing, and retrieval running directly on device.

**Building an AI agent that needs fast, private, scalable memory?** <ArrowLink href="https://cal.com/forms/5d3d4e31-ce22-4479-9e31-2b05051b35ef">Get in touch</ArrowLink>
]]></content:encoded>
            <author>Neha Varshneya</author>
            <category>case study</category>
            <category>Aside</category>
            <category>AI memory</category>
            <category>persistent memory</category>
            <category>AI browser</category>
            <category>on-device AI</category>
            <category>local retrieval</category>
            <category>semantic search</category>
            <category>embeddings</category>
            <category>AI agents</category>
            <enclosure url="https://mossdev.work/blog/aside-persistent-ai-memory/cover.jpg" length="0" type="image/jpg"/>
        </item>
        <item>
            <title><![CDATA[Why AI Infrastructure Is Moving Into the Runtime]]></title>
            <link>https://mossdev.work/blog/ai-infrastructure-moving-into-the-runtime</link>
            <guid isPermaLink="false">https://mossdev.work/blog/ai-infrastructure-moving-into-the-runtime</guid>
            <pubDate>Wed, 09 Sep 2026 09:00:00 GMT</pubDate>
            <description><![CDATA[Databases, analytics, caching, and compute all started as hosted services and ended up as libraries inside the application. Vector search and inference are crossing now, for the same three reasons: performance, cost, and developer experience. What moves into the runtime, what still belongs in the cloud, and why the cloud becomes the control plane.]]></description>
            <content:encoded><![CDATA[
Open the architecture diagram for a typical AI feature in 2026 and count the boxes. The application is one of them. Around it sit an embedding API, a hosted vector database, a reranking endpoint, and an LLM API, all in someone else's cloud, each billed separately and a network hop away from the code that needs it. Four extra services for a single feature, and for many teams, that architecture emerged incrementally rather than from a deliberate decision.

None of those services is a bad product, and the stack they form works. What makes it worth examining is that the rest of the developer stack used to look exactly like this, then stopped, and nobody misses the old shape. AI infrastructure has started down the same path, and the reasons change how production AI systems get built.

## Every Layer Eventually Becomes a Library

The most deployed database in the world ships as a library. SQLite is linked into every iPhone, every Android device, every major browser, and most operating systems, with nothing to connect to. Its win came from refusing the client-server model, at least for the enormous class of workloads that never needed one.

DuckDB repeated the move for analytics, running columnar queries inside your process over files already on your disk, at speeds that used to justify a warehouse. Once that worked, the warehouse round trip became optional for everything below true big-data scale.

The newest twist is the synced local copy. Turso's embedded replicas keep a full SQLite file inside your application and sync it against a remote primary, so reads happen locally in microseconds while the network only gets involved when synchronization requires it.

Caching followed the same curve from memcached clusters to in-process caches, and compute itself moved to the edge to sit next to the user, every layer heading the same way, into the runtime.

Hold that pattern up against AI infrastructure and most of the matches are already filled in.

| Capability | Hosted era | Runtime era |
| --- | --- | --- |
| Relational data | Database servers | SQLite linked into the app |
| Analytics | Cloud warehouses | DuckDB in-process |
| Caching | Managed cache clusters | In-memory caches inside the service |
| Compute | Centralized servers | Compute at the edge |
| Vector search | Hosted vector databases | In-process retrieval: FAISS, sqlite-vec, LanceDB |
| Inference | Frontier models behind APIs | Small models in OS runtimes and on device |

The bottom two rows are the ones being written right now, for the same three reasons the others crossed: performance, cost, and developer experience.

<figure className="art-figure art-figure--image reveal">
  <img src="/blog/ai-infrastructure-moving-into-the-runtime/absorption-arc.png" width="3200" height="1800" loading="lazy" alt="Each capability crossing from hosted service to runtime library. Relational data: database servers to SQLite in the app. Analytics: cloud warehouses to DuckDB in process. Caching: cache clusters to in-process cache. Compute: central servers to edge compute. Vector search: hosted vector databases to in-process retrieval with FAISS, sqlite-vec, and LanceDB, still mid-crossing. Inference: frontier APIs to on-device models, still mid-crossing." />
</figure>

## Performance: The Physics of the Network Hop

Reading from local memory takes around 100 nanoseconds, a round trip inside one datacenter costs around half a millisecond, and a round trip across a continent can cost around 150 milliseconds. Several orders of magnitude separate the first number from the last, and no amount of engineering on the far side of the socket can refund them.

Whether that matters depends on what you are building. A batch job never notices, but an interactive AI product lives inside those numbers. We covered what this does to retrieval in [The Retrieval Latency Tax](/blog/retrieval-latency-tax), so the short version here is that a call that looks fast on the provider's dashboard is still slower by the time it has crossed the network twice, and the crossing is the part users feel.

Latency also compounds in a way dashboards hide. An agent that makes five hosted calls per turn pays the sum of the hops rather than the average, and at P99 it pays the worst of each.

## Cost: Paying Retail for Compute You Already Own

Hosted AI services do the compute for you and charge for it with a margin. An embedding API and a managed vector database each run your workload on machines the vendor rents, and every price has to cover those machines plus the vendor's cut. Your own servers and your users' devices sit mostly idle while you pay for that second fleet.

A runtime moves some or all of that compute onto infrastructure you already have. You may still pay for storage, sync, and usage, but the heaviest part of the work, the per-query compute, now runs on hardware whose cost you were carrying anyway, which can make the total cheaper.

## Developer Experience: Setup Becomes a Package Install

The third argument is developer experience. Setting up a local package is usually easier than setting up infrastructure somewhere else, and libraries like SQLite and FAISS show how little it takes: one install command, one import, and your first working call. From there it's ordinary code inside your own project.

It also gives you back a laptop that works. An AI runtime that embeds, indexes, and queries in-process can behave the same in CI, on a plane, and in production, and the version you tested is the version you ship, pinned in your lockfile.

There is also one less thing that can go down. A hosted retrieval service can have an outage while your app is perfectly healthy, and your users still see a broken product. A library in your process has no separate status page: if your app is up, search is up.

## Retrieval Moved First

Vector search is the clearest case, because it has run through much of the arc in about three years. It began as a product category with dedicated hosted databases, and then the databases teams already ran absorbed it, natively in MongoDB, Redis, and SQL Server, and through the pgvector extension in Postgres. For many teams, vector search turned into a feature of the database they already had.

The next step is in progress: retrieval as a library. FAISS was always an in-process engine, sqlite-vec puts vector search in a single file with no daemon to run, and LanceDB is built for embedded and edge use. For many production workloads below tens of millions of vectors, retrieval can fit inside the process that needs it.

This is the step Moss is built for. A bare library hands you the index and leaves embedding, packaging, and keeping the data fresh as your problem, so we built the whole retrieval path to run in-process, embedding inference included. In our benchmark, a query over a 100,000-document index returns in 3.1 ms at P50 and 5.4 ms at P99. Numbers like that are what a query can cost once no network sits on the request path, with no exotic engineering required.

## The Operating System Is Now an AI Runtime

Inference is following, and the push is coming from an unexpected direction: the platforms themselves. Apple's Foundation Models framework makes an on-device model available to iOS apps through a system API, and supported workloads can run without a network round trip. Android exposes Gemini Nano to apps as an operating system service, Chrome ships a built-in model behind its Prompt API, and WebGPU inference in the browser is becoming increasingly viable.

The routing is the tell: on-device and cloud are becoming more interchangeable behind a runtime interface.

The developer side is keeping pace. Local inference tools have made running models on developer machines dramatically simpler, while model hubs now offer a large and growing selection of models packaged for local inference. The broader shift is toward matching the model to the task: many agent invocations are small, repetitive tasks that do not require a frontier model, making smaller models attractive on device and at the edge.

Edge AI used to describe an exotic deployment target, and increasingly it describes where inference can be cheapest and fastest to run.

## Real Time Is the Forcing Function

Cost and developer experience make the runtime attractive, and real-time products make it unavoidable. A voice agent has roughly 800 ms to start speaking before the pause reads as broken, and a stitched pipeline of hosted services can burn 50 to 200 ms of that budget on network hops alone, before any model has done any work. We broke the full turn down in [Building Voice AI That Feels Human](/blog/voice-ai-latency-budget).

The teams shipping voice agents, live copilots, and in-editor completion are moving infrastructure into the runtime because the latency budget leaves them nowhere else to put it.

## What Still Belongs in the Cloud

The shift has edges, and pretending otherwise would be selling something. Frontier-scale inference stays centralized, since the largest models need multi-node GPU clusters and the cloud economics of continuous batching and disaggregated serving are what make those tokens affordable. Very large vector indexes with heavy write churn can still benefit from a dedicated hosted engine, and an index living inside your process competes with your application for memory, which is a real cost on small devices.

There are operational edges too. Durability, backup, cross-device sync, and fleet-wide observability are things a runtime can't give itself, and when queries never touch your backend, knowing what your product is doing takes deliberate design.

The cloud keeps a large and permanent job while the request path moves out, workload by workload.

<figure className="art-figure art-figure--image reveal">
  <img src="/blog/ai-infrastructure-moving-into-the-runtime/deployment-runtimes.png" width="3200" height="1800" loading="lazy" alt="Where each layer runs. Browser: UI, session state, in-tab search and embeddings via WASM, for zero-network lookups where data never leaves the tab. Edge: token minting, routing, session bootstrap, light retrieval, close to users with fast cold starts. Device: on-device search, small models, offline agents, which work offline with deterministic latency and privacy by architecture. Cloud: frontier inference, index building, system of record, billing, for capability and coordination that need scale." />
</figure>

## The Cloud Becomes the Control Plane

The cloud becomes the control plane that trains and packages models, builds and stores indexes, syncs artifacts to wherever the application runs, meters usage, and coordinates fleets, while the runtime becomes the data plane that holds the working set and answers on the request path, in-process, in microseconds to low milliseconds.

It is the shape we build Moss around: the cloud builds and stores the index artifact, your process pulls it at load time and can poll for newer versions in the background, hot-swapping them with no query downtime, and every query runs inside your process. The control plane is what makes the runtime deployable at fleet scale.

<figure className="art-figure art-figure--image reveal">
  <img src="/blog/ai-infrastructure-moving-into-the-runtime/hosted-vs-runtime-stack.png" width="3200" height="1800" loading="lazy" alt="The hosted stack versus the runtime stack. Left: your app calls an embedding API, a vector database, a reranker, and an LLM API, each a 40 to 150 millisecond round trip. Right: your app contains the embedding model, the index, and a small model, with a single sync line to a control plane." />
</figure>

## The Diagram Gets Simpler

Each box on that opening architecture diagram survives the move. What changes is where it lives: inside your application, as code, with one quiet line back to the cloud for sync.

The database, the warehouse, and the cache all made this move before, and each time the diagram got simpler while the product got faster and cheaper. AI infrastructure is starting to follow the same path. The best AI infrastructure, like the best latency, is the kind that stops showing up in your diagram.

## Further Reading

- [The Production AI Stack: A Reference Architecture for Real-Time AI Systems](/blog/the-production-ai-stack)
- [Building Voice AI That Feels Human: A Latency Budget Breakdown](/blog/voice-ai-latency-budget)
- [The Retrieval Latency Tax: Why Your AI Agent Feels Slow](/blog/retrieval-latency-tax)
]]></content:encoded>
            <author>Sri Raghu Malireddi</author>
            <author>Grigory Tsyganok</author>
            <category>AI infrastructure</category>
            <category>AI runtime</category>
            <category>in-process retrieval</category>
            <category>vector search</category>
            <category>on-device AI</category>
            <category>edge AI</category>
            <category>embedded database</category>
            <category>latency</category>
            <category>real-time AI</category>
            <category>control plane</category>
            <enclosure url="https://mossdev.work/blog/ai-infrastructure-moving-into-the-runtime/og.png" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[Building Voice AI That Feels Human: A Latency Budget Breakdown]]></title>
            <link>https://mossdev.work/blog/voice-ai-latency-budget</link>
            <guid isPermaLink="false">https://mossdev.work/blog/voice-ai-latency-budget</guid>
            <pubDate>Wed, 19 Aug 2026 09:00:00 GMT</pubDate>
            <description><![CDATA[Voice AI has no typing indicator. Silence is the interface. A breakdown of the 800 millisecond time-to-first-audio budget across endpointing, speech recognition, retrieval, LLM inference, speech synthesis, transport, and playback, and how to keep the voice turn path small enough that users never say Hello twice.]]></description>
            <content:encoded><![CDATA[
The first sign that a voice agent is slow is usually not a latency graph. It is the user saying, "Hello?" a second time.

That second "Hello?" can create a surprisingly difficult failure loop. The agent may already be generating a response, but the user does not know that. They hear silence, assume the system did not hear them, and speak again. Speech recognition now has another utterance to reconcile, while the response that was already in flight may need to be interrupted. The user hears a clipped response, waits again, and may start changing how they speak to the system. Nothing is necessarily wrong with the underlying model; the problem is that the conversation has lost its timing.

This is one of the fundamental differences between Voice AI and chat applications. In a chat interface, latency can often be hidden behind visible progress. A typing indicator tells the user that the system is working, while streaming tokens allow them to start reading before the response is complete. Voice has no equivalent escape hatch. When the user finishes speaking, silence is the entire interface. They cannot see that a transcript is stabilizing, retrieval is running, or the model has started generating. If the response takes one or two seconds, the delay becomes ambiguous: did the agent hear me, is it still connected, or should I speak again?

That makes latency particularly important for Voice AI. A slow response does not simply make the application feel slower; it can interfere with turn-taking and cause the user and agent to speak over one another. The latency budget therefore has to be considered across the entire real-time AI stack, from endpointing and speech recognition through retrieval, inference, speech synthesis, transport, and playback. If one stage consumes more than its share, the effect is felt across the entire conversation.

## Why Voice AI Has Less Room Than Chat

Chat applications have several ways to communicate that work is happening. They can display a typing indicator, stream partial text, or allow users to scan a response as it is generated. These cues turn waiting into visible progress and give the user confidence that the system is still working.

Voice AI does not have that benefit. Once the user yields the floor, the system has to communicate progress through its response. A period of silence provides no information about what is happening internally. This is why a delay that feels insignificant in a chat application can feel much more pronounced in a voice conversation.

The problem also becomes more serious as latency increases because the user can change their behavior. A user who waits one second may simply notice the delay. A user who waits two seconds may start wondering whether the system heard them. A user who waits several seconds may repeat the question. That second utterance can arrive while the first response is already being generated, forcing the system to detect the interruption, cancel work that is no longer relevant, and recover the conversational state.

In other words, Voice AI latency is not simply a measure of how quickly a system produces an answer. It affects the system's ability to correctly determine who has the floor.

<figure className="art-figure art-figure--image reveal">
  <img src="/blog/voice-ai-latency-budget/latency-budgets.png" width="3200" height="1800" loading="lazy" alt="Illustrative latency references by product. Chat assistants target about one second to first token, while Voice AI uses about 800 milliseconds to first audio as a design reference because silence provides no visible progress cue." />
</figure>

## The Latency Budget for a Single Voice AI Turn

For a voice agent, one of the most useful end-to-end metrics is **time to first audio (TTFA)**:

```
TTFA = first_audio_played_at_user - end_of_user_speech
```

This measures the time between the user finishing their utterance and actually hearing the agent begin its response. That distinction matters because individual provider metrics do not necessarily represent the experience the user is having. A speech recognition provider returning a transcript quickly, an LLM producing its first token, or a text-to-speech service returning its first byte are all useful measurements for diagnosing the system, but none of them tells you when the user actually hears the response.

Our [Production AI Stack](/blog/the-production-ai-stack) uses approximately **800 milliseconds to first audio as a Voice AI design reference**. This should be treated as a design target rather than a universal benchmark or customer SLA. The appropriate target will vary depending on the type of conversation, language, network conditions, device, and acceptable error rate. The important thing is to establish an explicit budget before making architectural decisions.

A representative turn can be broken down as follows:

| Stage | What must happen | Illustrative reference |
| --- | --- | --- |
| Endpointing | Determine that the user has yielded the floor | Part of ~175 ms of shared residual headroom |
| Final speech transcript | Stabilize the words needed to route the turn | ~100 ms best case, ~200 ms typical |
| Retrieval | Fetch context for a grounded answer | ~250 ms best case, ~500 ms to 1.5 s typical |
| LLM | Produce the first useful tokens | ~200 ms best case, ~400 ms typical |
| Speech synthesis | Turn a stable phrase into the first audio | ~75 ms best case, ~200 ms typical |
| Transport and playback | Deliver, buffer, decode, and play audio | Part of ~175 ms of shared residual headroom |
| **Full turn** | **User stops speaking to first audio played** | **800 ms design reference** |

The four quantified stages consume approximately 625 milliseconds in the best case. Against an 800-millisecond design reference, that leaves roughly 175 milliseconds for endpointing, transport, playback, and orchestration. In the typical ranges, those same four stages add up to approximately 1.3 to 2.3 seconds before separately measured endpointing and playback time are included.

This is why latency cannot be optimized one component at a time. A retrieval call that takes an additional 200 milliseconds, for example, may not look particularly concerning in isolation. But when that delay comes out of the same budget as inference, speech synthesis, and playback, it can be the difference between a response that feels immediate and one that causes the user to wonder whether the agent heard them.

## Following a Voice AI Request Through the Stack

Consider a user asking, "Can I use this in a mobile app?" and then stopping speaking. The endpointing system first needs to determine that the user has actually finished the turn rather than simply pausing in the middle of a sentence. Speech recognition needs to stabilize the final transcript so that the application can correctly route the request. The application then retrieves the relevant product context, passes that context to the language model, and begins generating a grounded response. Once enough stable language is available, speech synthesis can begin producing audio, which then has to travel back to the device, be buffered and decoded, and ultimately be played through the speaker.

Some of these stages can overlap when doing so does not compromise correctness. For example, partial transcripts can be used to begin speculative retrieval, and stable portions of model output can be streamed directly into speech synthesis. However, the underlying dependencies remain. Retrieval needs a sufficiently stable query, generation needs the retrieved context, speech synthesis needs enough stable language to form a coherent phrase, and playback needs audio to have reached the client.

The user does not experience these as separate operations. They experience the sum of them as the amount of time between finishing a sentence and hearing a response.

That makes each stage worth examining individually.

## 1. Endpointing and Speech Recognition: Buy Certainty Deliberately

The first latency decision happens before retrieval or inference. The system has to determine whether a pause means "I am done" or "I am still thinking."

A fixed silence timer is easy to implement, but making it consistently natural is difficult. If the system waits too long, every response begins with unnecessary dead air. If it commits too quickly, it can cut users off before they finish dates, product names, email addresses, corrections, or multi-clause questions. Reducing endpointing latency without measuring false endpoints therefore risks replacing one poor experience with another.

Endpointing should be evaluated using paired performance metrics, including end-of-turn decision delay, false endpoint rate, transcript correction after the endpoint, and results segmented by intent, language, device, and network conditions. A short confirmation may support aggressive endpointing, while a technical question or email address may contain pauses that should remain part of the same turn.

Partial transcripts can also be useful for speculative routing or retrieval, provided that the work is inexpensive to cancel. The final transcript should remain authoritative whenever the user's last words change the meaning of the request. This allows the system to use streaming information to save time without sacrificing correctness.

The objective is therefore not simply to detect silence as quickly as possible. It is to determine, with sufficient confidence, when the user has actually yielded the floor.

## 2. Retrieval: Protect the Middle of the Turn

Retrieval is one of the places where an otherwise fast Voice AI system can lose a significant portion of its latency budget.

A hosted retrieval request is more than a search operation. Depending on the architecture, it can involve connection acquisition, authentication, load balancing, queueing, network transit, query execution, serialization, and deserialization. The [Production AI Stack](/blog/the-production-ai-stack) describes this as a retrieval latency tax: infrastructure overhead incurred before the model receives the context it needs. If an agent makes multiple dependent retrieval requests, that overhead can be paid multiple times within a single turn.

In a chat application, several hundred milliseconds of retrieval latency may be difficult for a user to notice. In Voice AI, the same delay becomes part of the silence between turns. When grounding consistently takes hundreds of milliseconds, teams may respond by skipping retrieval for requests that appear simple or by putting more information directly into the prompt. The first approach increases the risk of ungrounded answers, while the second can make prompts larger and less selective.

An alternative is to keep an already-loaded index in the agent runtime so that retrieval does not require a network request on every turn.

In the benchmarks published with the Production AI Stack, moving retrieval in process reduced median latency from 67 milliseconds to 5 milliseconds and P99 latency from 222 milliseconds to 13.5 milliseconds. In a published 100,000-document benchmark at top-k five, including embedding inference, Moss measured 3.1 milliseconds at P50 and 5.4 milliseconds at P99. These are published benchmark figures rather than guarantees for every corpus, machine, filter, or concurrency level, so teams should measure their own workloads.

The broader lesson is that retrieval performance is not just about making search faster. Lower and more predictable retrieval latency creates additional headroom for every other part of the Voice AI pipeline.

<figure className="art-figure art-figure--image reveal">
  <img src="/blog/voice-ai-latency-budget/benchmark-query-latency.png" width="3200" height="1800" loading="lazy" alt="Published query latency benchmark over a 100,000 document index at top k five, including embedding inference. Moss in-process retrieval measures 3.1 milliseconds at P50 and 5.4 milliseconds at P99, compared with hosted retrieval systems that include network service overhead." />
</figure>

## 3. LLM Inference and TTS: Optimize the First Speakable Phrase

Time to first token is an important LLM metric, but it is not the same as the time at which a voice agent can actually begin responding.

The first token may be punctuation, an incomplete fragment, or language that is not yet stable enough to synthesize naturally. For Voice AI, there are therefore two connected budgets: the time required for the model to produce useful, stable language and the time required for speech synthesis to turn that language into playable audio.

Streaming between these stages is critical. Rather than waiting for the entire model response, the system can pass sufficiently stable output into speech synthesis and begin producing audio while the rest of the response is still being generated. However, chunk size needs to be tuned carefully. Very small chunks can reduce apparent latency while producing unnatural or fragmented speech, while large chunks introduce a buffer that effectively recreates the latency the system was trying to eliminate.

For this reason, teams should measure time to first token, time to first stable speakable phrase, and time to first audio separately. A model dashboard can report excellent time-to-first-token performance while the user is still waiting because the application has not yet produced enough stable language to speak.

## 4. Transport and Playback: Stop the Clock at the Speaker

A speech synthesis provider returning audio is not the same thing as the user hearing that audio. The response still needs to travel across the network, reach the client, enter a playback buffer, be decoded, and begin playing without immediately stalling.

This final stage is easy to overlook because most infrastructure dashboards stop measuring before the user experience actually begins. Voice AI systems should instead measure client receipt, buffer time, decode time, and first playback, along with the region and network conditions affecting those measurements.

Keeping unavoidable services in compatible regions and reusing connections can help reduce transport overhead, but the most important architectural decision is simply to measure the complete path. Provider-level timings explain individual components; microphone-to-speaker timing explains the experience.

## Keep the Voice AI Turn Path Small

Many of the biggest latency improvements come not from making every component marginally faster, but from removing unnecessary work from the repeated turn path altogether.

Static assets should be built and validated before serving traffic. Reusable clients, connections, and models can be prewarmed at worker startup. Tenant configuration, knowledge, and conversation state can be loaded at session start rather than repeatedly on every turn. The critical path can then remain focused on the work that actually needs to happen for each interaction: endpointing, turn-specific retrieval, generation, synthesis, and playback.

This distinction is particularly important for first-turn latency. Cold-start work hidden inside the first request is still user-facing latency, even if the system looks fast once it is warm. First-turn and warm-turn distributions should therefore be reported separately rather than allowing a healthy steady-state median to conceal a poor first impression.

Streaming also needs to extend across every boundary. A single batch-oriented stage can erase the gains produced by streaming everywhere else in the stack.

There is a second critical path to consider as well: interruption handling. While the agent is speaking, the system must detect meaningful user speech, stop playback, cancel generation and synthesis, and preserve only the words the user actually heard. Interruption stop time should therefore be measured alongside false interruption and false barge-in rates. A system that responds quickly but treats every cough or background noise as a new turn is not necessarily delivering a better conversational experience.

## Founding Agent as a Practical Example

Moss's [Founding Agent](/blog/founding-agent), a voice landing-page assistant, applies these same latency principles in a real product experience.

Streaming speech recognition handles the turn decision, while deterministic navigation and scroll intents can bypass a model call altogether. Substantive questions retrieve context from an in-process Moss index, and stable language streams into speech synthesis. Tenant configuration and the knowledge index are loaded at session lifetime rather than inside every individual turn.

Each of these choices protects a different part of the latency budget. Some allow work to begin at the right moment, some eliminate unnecessary work, some remove network hops, and others keep audio moving through the pipeline. Together, they demonstrate why Voice AI latency is fundamentally an architectural problem rather than a single-provider optimization problem.

## Measure the Conversation, Not the Provider Dashboard

A useful Voice AI trace should follow each turn from microphone to speaker using a shared identifier. At minimum, the trace should include endpoint decision time, final transcript time, route selection, retrieval wall time, prompt assembly, model time to first token, time to first stable speakable phrase, speech synthesis first byte, client receipt, buffer time, first playback, and cancellation timestamps for interrupted turns.

Each stage should have an owner, a measurement boundary, and a tail-latency objective. P50, P95, and P99 should be aggregated by region, device class, network type, and turn number. First-turn latency should be separated from warm-turn latency, model time to first token from first speakable phrase, and synthesis first byte from first audio actually played.

Those measurements also need to be considered alongside conversational quality. False endpoint rate, false barge-in rate, transcript correction rate, first-turn cold-start rate, and the percentage of substantive answers that were grounded all affect whether the user perceives the system as working well.

A latency improvement that produces more interruptions or fewer grounded answers is not necessarily an improvement. It may simply have moved the problem somewhere else in the stack.

## The Best Latency Is Invisible

There is no single provider that can make a Voice AI system feel conversational on its own. A faster language model cannot recover a slow endpoint decision. Fast speech synthesis cannot speak context that is still crossing a retrieval boundary. And a low median latency cannot compensate for repeated tail-latency stalls.

The entire stack has to work within one small budget.

That means defining time to first audio as an end-to-end metric, assigning every stage a share of the budget, streaming across boundaries, moving repeated setup out of the turn, and removing unnecessary network work from the critical path. It also means measuring latency alongside the correctness of turn-taking, because speed without conversational accuracy does not produce a better voice experience.

When those pieces work together, the user asks a question, yields the floor, and hears a grounded answer when they expect one. The system does not need to explain that it is still working because the response arrives naturally enough that the user never has to wonder.

There is no second "Hello?" because there is no gap to explain.

The best Voice AI latency is the latency the user never notices.

## Further Reading

- [The Production AI Stack: A Reference Architecture for Real-Time AI Systems](/blog/the-production-ai-stack)
- [The Retrieval Latency Tax: Why Your AI Agent Feels Slow](/blog/retrieval-latency-tax)
- [We Built a Voice AI Agent for Our Website](/blog/founding-agent)
]]></content:encoded>
            <author>Sri Raghu Malireddi</author>
            <author>Abhishake Kumar Bojja</author>
            <category>voice AI</category>
            <category>latency budget</category>
            <category>time to first audio</category>
            <category>TTFA</category>
            <category>endpointing</category>
            <category>speech recognition</category>
            <category>retrieval latency</category>
            <category>LLM inference</category>
            <category>text to speech</category>
            <category>real-time AI</category>
            <category>conversational AI</category>
            <enclosure url="https://mossdev.work/blog/voice-ai-latency-budget/og.png" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[The Production AI Stack: A Reference Architecture for Real-Time AI Systems]]></title>
            <link>https://mossdev.work/blog/the-production-ai-stack</link>
            <guid isPermaLink="false">https://mossdev.work/blog/the-production-ai-stack</guid>
            <pubDate>Sat, 18 Jul 2026 09:00:00 GMT</pubDate>
            <description><![CDATA[The seven layers of every real-time AI system - models, inference, search, memory, sessions, orchestration, and deployment - and where latency comes from at each one. A reference architecture for building voice agents, copilots, and conversational AI that feel instant.]]></description>
            <content:encoded><![CDATA[
A founder built an AI support agent for his ecommerce store using what has become the standard modern AI stack: a large language model, a cloud vector database containing the company's help center, and a small set of tools for common actions like looking up orders and initiating returns. The system worked well. Customers could ask questions, check the status of an order, or begin a return without ever needing to speak with a human.

The problem wasn't correctness. It was latency.

Every interaction followed the same execution path. A customer asked a question, the agent queried the vector database for relevant context, waited several hundred milliseconds for the results to return, and only then could the model begin generating a response. Nothing in the architecture was technically broken, but the cumulative delay was obvious enough that conversations felt slower than they should.

Like many teams building production AI systems, he optimized for responsiveness by removing retrieval altogether.

Instead of retrieving context on every turn, he embedded everything directly into the system prompt. The help center, shipping policies, return rules, pricing information, FAQs, and any other information a customer might reasonably ask about all became part of the prompt that accompanied every request. It eliminated an external dependency, reduced latency, and simplified the architecture.

For a while, it looked like the right tradeoff.

As conversations became longer, however, the architecture started failing in more subtle ways. Every message added another layer of conversational history to a context window that was already carrying the company's knowledge base, forcing the model to reason over an increasingly large body of information. As the available context filled up, retrieval accuracy was effectively replaced by probabilistic recall. Shipping questions began pulling in return policies, discounts from unrelated products appeared in responses, and facts that had been clear at the beginning of the conversation became less reliable over time.

The most frustrating part was that these failures rarely appeared during testing. Short benchmark prompts continued to perform well, while the conversations that actually mattered in production were the ones that exposed the weaknesses of the architecture. Customers asked follow up questions, referred back to earlier messages, changed their minds halfway through a workflow, and expected the system to maintain consistency across dozens of conversational turns. Those were precisely the interactions where the model became least reliable.

This pattern appears across every category of production AI application, from customer support agents and enterprise copilots to voice AI platforms and conversational search systems. Teams often assume they're dealing with a prompting problem or a model quality problem, when in reality they're running into architectural limits. Large prompts eventually become expensive to process, context windows inevitably become saturated, and relying on a language model to remember everything produces systems that become less predictable as conversations grow longer.

## The Production AI Stack

Every production AI system, whether it's powering a voice AI platform, an enterprise copilot, or a conversational search experience, is ultimately composed of the same seven architectural layers:

1. Models determine which foundation model is responsible for each task.
2. Inference controls where and how token generation happens.
3. Search retrieves the knowledge required to answer each request.
4. Memory persists information about users and previous interactions.
5. Sessions maintain conversational state across multiple requests.
6. Orchestration coordinates the execution of every component throughout the conversation.
7. Deployment determines where each layer runs, whether in the browser, at the edge, on device, or in the cloud.

Each of these layers can be optimized independently, but production AI systems are rarely limited by any single component. Performance, latency, reliability, and cost emerge from the interactions between them. Understanding those tradeoffs requires looking at the entire architecture rather than any individual layer, which is where we'll begin.

## Start With the Latency Budget

Every production AI system begins with the same constraint: latency.

Before deciding which model to use, how to structure retrieval, or where to run inference, you need to understand how much latency your users will actually tolerate. Every interactive product has a latency budget, and once that budget is exceeded, no amount of model quality can recover the user experience.

Human computer interaction research established these thresholds decades before large language models existed. Responses under roughly 100 milliseconds feel instantaneous. Around one second, users remain in their flow of thought but become aware that they're waiting. Beyond ten seconds, attention shifts elsewhere. IBM's work on the Doherty Threshold arrived at a similar conclusion in the early 1980s, arguing that interactive systems should respond within roughly 400 milliseconds to maintain a continuous feedback loop between the user and the computer.

Conversation is even less forgiving. Human turn taking typically happens within about 300 milliseconds, which means every millisecond an AI system spends retrieving data, waiting on network calls, or generating tokens directly competes with a rhythm that people have spent their entire lives expecting.

The acceptable latency budget therefore depends on the product you're building:

<figure className="art-figure art-figure--image reveal">
  <img src="/blog/the-production-ai-stack/latency-budgets.png" width="3200" height="1800" loading="lazy" alt="Latency budgets by product. Inline code completions: 100 to 400 ms, because suggestions must arrive before the user continues typing. Chat assistants: about 1 second to first token, so responses preserve the user's train of thought. Voice AI agents: about 800 ms to first audio, so conversations feel natural and interruptible. Background agents: minutes to hours, where latency compounds into overall completion time and infrastructure cost." />
</figure>

These aren't arbitrary targets. They directly influence user behavior.

Google demonstrated this in a controlled experiment in 2009 by intentionally adding between 100 and 400 milliseconds of server side latency to search results for a randomized group of users. Even delays at the lower end of that range caused people to perform fewer searches, with engagement steadily decreasing as latency increased. Small delays that seem insignificant in isolation become measurable product problems when they occur millions of times every day.

Voice AI makes these constraints impossible to ignore.

Among all AI applications, voice has the smallest latency budget while requiring the largest number of sequential operations before a response can begin. A typical conversational turn requires speech recognition, retrieval, language model inference, and speech synthesis to happen one after another. Because each stage depends on the output of the previous one, their latencies accumulate rather than overlap.

A representative production pipeline looks something like this:

<figure className="art-figure art-figure--image reveal">
  <img src="/blog/the-production-ai-stack/voice-turn-latency.png" width="3200" height="1800" loading="lazy" alt="A representative voice pipeline, best case versus typical production. Speech to text final transcript: 100 ms best, 200 ms typical. Retrieval from hosted vector database: 250 ms best, 500 ms to 1.5 s typical. LLM time to first token: 200 ms best, 400 ms typical. Text to speech first audio: 75 ms best, 200 ms typical. Total time to first response: about 625 ms best, 1.3 to 2.3 s typical." />
</figure>

One number stands out immediately.

The language model isn't usually the largest source of latency. Retrieval is.

Despite often being treated as solved infrastructure, retrieval is frequently both the slowest and most variable stage in the request pipeline. Under ideal conditions it consumes a substantial portion of the latency budget, and under realistic production conditions it often exceeds the entire budget before the model has generated a single token.

The obvious question is why.

## Why Web Infrastructure Breaks Down for Real Time AI

The default AI stack inherited its architecture from the web.

For more than a decade, web applications have followed the same pattern: keep application servers stateless, centralize data in managed services, and connect everything over the network. That architecture is highly effective for traditional web workloads because most user interactions involve only a handful of network requests, and adding another database query rarely changes the user experience in any meaningful way.

Production AI systems have a fundamentally different execution model.

Instead of serving a page with one or two database lookups, an AI application performs retrieval, tool execution, model inference, memory access, reranking, and additional retrieval repeatedly throughout a conversation. Every conversational turn becomes a pipeline of dependent operations, each introducing another network boundary and another opportunity for latency to accumulate.

The industry is already experiencing the consequences. LangChain's State of Agent Engineering survey found that among teams already deploying agents in production, latency ranks as the second largest engineering challenge, surpassed only by output quality.

The reason becomes obvious once you look beyond raw network latency.

A network hop isn't simply the time required to transmit packets between machines. Every request also incurs serialization, TLS negotiation, authentication, connection pooling, load balancing, scheduling, and queueing before the remote service begins executing. Even within a single cloud region, that overhead commonly adds 40 to 150 milliseconds to every request. Cross region communication can easily increase that to 150 to 400 milliseconds, all before the model has processed a single token.

Those costs become even more significant because AI systems rarely perform one isolated request.

Google's paper The Tail at Scale showed that once requests fan out across multiple services, infrequent slow responses begin to dominate overall system latency. Production AI systems exhibit the same behavior, except instead of making many requests in parallel, they often perform them sequentially. The tail latency of one service becomes the starting point for the next, causing delays to compound across an entire conversational turn.

This is why median latency is often a misleading metric.

A managed vector database might appear perfectly acceptable at P50 while exhibiting dramatically higher latency at P99 because of cold connections, garbage collection pauses, noisy neighbors, or temporary resource contention. In our own production benchmarks, cloud vector database round trips commonly cluster between 500 and 900 milliseconds at P99.

Users never experience your median latency.

They experience your worst moments.

In a conversation lasting twenty turns, those worst case requests stop being statistical outliers. They become inevitable. And once retrieval consumes half a second or more before inference even begins, no language model can make the system feel responsive.

## Layer 1: Models

Models are the layer most teams spend the most time debating, yet they're rarely the primary bottleneck in a production AI system. Foundation models have become remarkably capable, latency continues to improve, and switching between providers is easier than ever. The challenge is no longer finding the best model. It's using the right model for the right job.

Production AI systems shouldn't be built around a single model. They should be built around a portfolio of models optimized for different workloads. Frontier models are reserved for tasks where deeper reasoning justifies the additional latency and cost. Faster models handle routing, classification, tool selection, and other decisions that occur on nearly every turn. Smaller or local models power embeddings, reranking, moderation, and guardrail checks. Most requests in a production system are relatively simple, and routing those decisions to a model that responds in a few hundred milliseconds instead of more than a second is often the easiest latency improvement you'll make.

For interactive applications, time to first token (TTFT) is a far more meaningful metric than benchmark scores or tokens per second. Users notice how quickly a response begins, not how quickly it finishes. TTFT varies significantly across hosted models and grows with prompt size, which means published benchmarks are rarely representative of production workloads. Measure it yourself using realistic prompts and conversation histories rather than synthetic benchmarks.

Model configuration also has a direct impact on latency. Reasoning modes, thinking budgets, and other advanced inference settings can add seconds to every request. Those tradeoffs may be worthwhile for complex planning or analysis, but for most grounded conversational interactions they introduce noticeable delays while providing little measurable improvement in response quality.

Finally, assume every model you choose today will eventually be replaced. The pace of model releases makes portability a practical engineering requirement rather than an architectural ideal. Keeping prompts in version control, maintaining a reliable evaluation suite, and abstracting model providers behind a consistent interface makes it possible to adopt better models as they become available without rebuilding the rest of your system.

## Layer 2: Inference

If models determine what generates the response, inference determines where and how that response is generated. Those decisions establish both the latency floor and the long term cost profile of your system.

The first rule is simple: stream everything. Production AI systems should never wait for an entire completion before responding. Measure time to first token and tokens per second independently because they optimize for different outcomes. TTFT determines perceived responsiveness, while generation speed influences throughput, infrastructure utilization, and overall cost.

Not every response needs to come from a language model. Greetings, acknowledgments, confirmations, and other predictable conversational patterns can often be served from templates, and in voice applications they can be sent directly to text to speech while the model continues reasoning in the background. Delivering the first audio or visual feedback a few hundred milliseconds earlier often makes the entire interaction feel dramatically faster, even when the underlying inference time remains unchanged.

Inference also benefits from aggressive caching. System prompts, tool definitions, and other static prefixes rarely change between requests, making them ideal candidates for prefix caching. Most hosted inference providers and self hosted serving frameworks support this optimization, allowing repeated prompt tokens to be reused instead of reprocessed. For applications with large system prompts, it's one of the simplest and highest impact ways to reduce TTFT.

Finally, choose your inference environment intentionally rather than by default. Hosted APIs provide access to the latest frontier models but introduce network latency and queue variability. Self hosted GPUs offer predictable performance and greater operational control, but require significant infrastructure investment. On device inference eliminates network latency entirely while improving privacy, although current models remain more constrained. Increasingly, production AI systems combine all three, selecting the execution environment that best matches the latency, capability, and cost requirements of each request.

## Layer 3: Search

Every production AI system retrieves information. Whether it's querying a knowledge base, product catalog, documentation, user history, or internal company data, retrieval sits on the critical path of nearly every grounded response. The default architecture has been to attach a hosted vector database to the stack and call it Retrieval Augmented Generation (RAG). That decision has become so common that it's rarely questioned, even though it's often the single largest source of user facing latency.

We refer to this as the retrieval latency tax: the 100 to 500 milliseconds of infrastructure overhead incurred before the model has even seen the context it needs to answer the question. Most of that time isn't spent searching. It's spent crossing the network.

Consider what actually happens during a retrieval request. A customer asks where an order is, the application generates an embedding, sends it to a hosted vector database, waits for the top results to come back, and only then constructs the prompt for the language model. The database itself typically spends only a few milliseconds performing the nearest neighbor search. Everything else is serialization, authentication, network transit, load balancing, and queueing.

The solution isn't simply choosing a faster vector database. It's removing the network from the hot path entirely.

When retrieval executes inside the same runtime as the agent, search starts behaving like a local function call instead of a distributed systems problem. In our own profiling, moving retrieval from a hosted service into the application process reduced median latency from 67 milliseconds to 5 milliseconds, while P99 latency dropped from 222 milliseconds to 13.5 milliseconds. The improvement wasn't the result of a better indexing algorithm. It came from eliminating unnecessary infrastructure between the application and its data.

Our published benchmark illustrates the same pattern across a 100,000 document index with a top k of five, including embedding inference:

<figure className="art-figure art-figure--image reveal">
  <img src="/blog/the-production-ai-stack/benchmark-query-latency.png" width="3200" height="1800" loading="lazy" alt="Query latency across a 100,000 document index, top k of five, including embedding inference. Moss in process: 3.1 ms P50, 4.3 ms P95, 5.4 ms P99. ChromaDB: 351.8 ms P50, 423.5 ms P95, 538.5 ms P99. Pinecone: 432.6 ms P50, 732.1 ms P95, 934.2 ms P99. Qdrant: 597.6 ms P50, 682.0 ms P95, 771.4 ms P99." />
</figure>

The difference isn't search quality. It's architecture. Hosted databases spend most of their latency budget moving requests across infrastructure. In process retrieval spends nearly all of it searching.

Several concerns naturally follow. The first is memory usage. Fortunately, most production knowledge bases are far smaller than many teams assume. A corpus of 10,000 to 100,000 chunks typically requires tens to hundreds of megabytes of memory depending on the embedding dimensions and quantization strategy, making it practical to run inside a server process, browser, or even a modern mobile device.

The second concern is maintaining a source of truth. Local retrieval doesn't replace centralized infrastructure. Instead, it separates the control plane from the data plane. The cloud remains responsible for building, versioning, and distributing indexes, while the runtime executes queries against a local copy. Updates arrive through background synchronization and hot swapped indexes rather than synchronous requests on every conversational turn, following the same architectural pattern that CDNs introduced for web applications years ago.

Finally, semantic search alone is rarely sufficient. Product names, SKUs, error codes, and other exact identifiers don't embed particularly well. Production retrieval systems therefore combine semantic similarity with keyword search, blending both signals into a single query so they can retrieve conceptual matches and exact matches without sacrificing latency.

## Layer 4: Memory

Agent memory often feels like a fundamentally new capability, but in practice it's another retrieval problem.

Every memory operation ultimately answers the same kinds of questions. What has this user told us before? What happened earlier in this conversation? Which preferences should carry over into future sessions? Each of those is simply a search over accumulated state.

That means memory inherits the same architectural constraints as retrieval. If conversation history lives behind another hosted service, every conversational turn introduces an additional network request before the model can respond. The system pays the retrieval latency tax once to retrieve external knowledge and again to remember what it already knows about the user.

It helps to distinguish the three different kinds of memory because each has a different lifetime and belongs in a different part of the architecture.

Working context consists of the messages, retrieved documents, and tool outputs included in the current prompt. It exists only for the duration of a single model invocation.

Session memory captures information established throughout an ongoing conversation, such as user preferences, extracted facts, unresolved tasks, and conversation summaries. Because it's queried continuously, it needs to be available at conversational speed.

Long term memory persists across sessions. It represents durable information about the user, changes relatively infrequently, and is read much more often than it is written.

Production systems generally reflect those different lifetimes in their implementation. Session memory remains inside the runtime as a small, mutable local index that is continuously updated as the conversation evolves. Rather than storing raw transcripts, the agent writes distilled facts, summaries, and structured state, improving retrieval quality while keeping the index compact. Long term memory, meanwhile, is synchronized asynchronously with a persistent system of record after responses are sent or sessions end. When the next conversation begins, that memory is loaded once into the runtime and queried locally alongside the knowledge base, avoiding another network dependency on every turn.

## Layer 5: Sessions

A session is the unit of state in a conversational AI system. One user, one conversation, one evolving context. Unlike traditional web applications, where requests are largely independent and can be routed to any server, AI systems maintain state over minutes or even hours while multiple components update that state simultaneously. The language model, tool calls, user input, and memory layer are all contributing to the same conversation, making session management a core architectural concern rather than an implementation detail.

Most production issues fall into one of two categories.

The first is state leaking across sessions. If two conversations can ever reference the same mutable state, eventually one user will inherit another user's context. These bugs are often rare, difficult to reproduce, and severe when they occur. Every session architecture therefore needs a clear isolation boundary, whether that's a dedicated process, runtime, or another execution model. The important property isn't the implementation itself but the guarantee that state cannot cross that boundary, even under failure conditions.

The second failure mode is state fragmentation within a session. The model maintains its own conversational context, each tool often tracks its own internal state, and the client may keep yet another representation. Over time those views inevitably diverge. An agent might confirm a reservation that has already been canceled because one tool consulted stale state while another had already processed the update. The solution is to establish a single authoritative representation of the session and treat every other view, including the model's context window, as a derived cache rather than the source of truth.

Latency introduces a third challenge. Many of the costs associated with starting a conversation are paid on the very first request, exactly when users are forming their initial impression of the system. A useful way to reason about initialization is to classify every operation by how frequently it should occur: once per deployment, once per worker, once per session, or once per conversational turn. Model loading belongs at worker startup, not session creation. User state should be fetched when the session begins, not when the first question arrives. Ideally, the first turn performs only the work required to answer the first turn.

Finally, sessions must survive failure. Network connections drop, browser tabs close, workers restart, and users frequently return after long periods of inactivity. Persisting session snapshots allows conversations to resume instead of restarting from scratch, while recognizing that capacity planning shifts from requests per second to concurrent active conversations. In production, you're no longer sizing infrastructure around HTTP requests. You're sizing it around live sessions.

## Layer 6: Orchestration

Orchestration coordinates the execution of every layer in the stack. It receives user input, decides what happens next, invokes tools, retrieves context, calls models, updates memory, and repeats that process for every conversational turn. Frameworks such as LiveKit Agents, Pipecat, LangChain, and the Vercel AI SDK provide abstractions for building these execution loops, but the orchestration layer itself is rarely the limiting factor. As Anthropic has observed, the most successful production systems typically rely on simple, composable patterns rather than increasingly complex orchestration frameworks.

What orchestration does amplify is infrastructure latency.

Every tool invocation inherits the cost of the service behind it. An agent that performs three hosted lookups during a single conversational turn doesn't pay the latency once. It pays it three times. A 300 millisecond network request quickly becomes nearly a full second of waiting before the model can produce a response. Replace those same operations with in process function calls and the orchestration logic remains identical, but the user experiences an entirely different system.

That leads to a useful design principle: everything that executes on the hot path should behave like a function call. Network requests should either happen infrequently or be moved off the critical path altogether. Invoking a booking API once during a conversation is rarely a problem. Querying a remote knowledge service on every conversational turn is.

Grounding illustrates why this distinction matters. The most reliable production agents retrieve fresh context before answering substantive questions rather than relying on the model's parametric memory. That strategy significantly improves factual accuracy, but only if retrieval is inexpensive enough to perform on every turn. When retrieval consistently costs several hundred milliseconds, teams begin skipping it to reduce latency. When retrieval behaves like a local function call, grounding becomes the default rather than an optimization that must be rationed.

## Layer 7: Deployment

The final layer is deployment, although it's often the first question teams ask: Where should the application run?

That's the wrong question.

The better question is where should each layer run? A production AI system is no longer deployed to a single environment. Different parts of the stack belong in different places depending on their latency, privacy, and compute requirements.

<figure className="art-figure art-figure--image reveal">
  <img src="/blog/the-production-ai-stack/deployment-runtimes.png" width="3200" height="1800" loading="lazy" alt="Where each layer runs. Browser: UI, session state, in tab retrieval, embeddings via WASM, for zero network lookups and private by default. Edge: authentication, routing, session bootstrap, lightweight retrieval, close to users with fast startup times. Device: local retrieval, small language models, offline agents, for deterministic latency, offline support, and privacy. Cloud: frontier models, index construction, synchronization, billing, system of record, for centralized coordination and scalable compute." />
</figure>

The same architectural principle that applies to retrieval applies to the entire stack: the cloud should operate as the control plane, while the browser, edge, and device form the data plane. The control plane is responsible for building indexes, distributing updates, synchronizing state, and coordinating infrastructure. The data plane executes the hot path as close to the user as possible.

This represents a meaningful shift from the architecture that defined the last decade of web infrastructure. Traditional applications centralized both state and compute behind managed services because capable computation could only happen inside a data center. AI systems inherited that architecture by default. That assumption is becoming less true every month.

Foundation models are rapidly commoditizing. Performance differences between leading providers continue to narrow, switching costs continue to fall, and the model itself is becoming one of the most replaceable components in the stack. At the same time, local capabilities continue to improve. Small language models are increasingly capable, retrieval already runs comfortably on consumer hardware, and modern browsers, laptops, and mobile devices now include hardware accelerators designed specifically for AI workloads.

As those trends converge, the default architecture shifts toward local first execution. The cloud remains essential for coordination, synchronization, and heavyweight inference, but an increasing portion of the user experience no longer needs to depend on a network request.

## If I Were Building a Production AI System Today

Every architectural decision would start with a single constraint: the latency budget. Rather than selecting technologies first and measuring performance later, I'd define the target latency for the product and work backward from there. Every component in the stack would need to justify the latency it adds.

The resulting architecture would look like this:

| Layer | Recommendation | Why |
|---|---|---|
| Models | Route across multiple models. Use frontier models for reasoning, fast models for routing and tool selection, and small local models for embeddings, reranking, and guardrails. | Most requests don't require the largest model. |
| Inference | Stream every response, enable prefix caching, and avoid unnecessary model calls with templates where appropriate. | Minimize time to first token and improve perceived responsiveness. |
| Search | Run hybrid semantic and keyword retrieval in process. Build and distribute indexes from the cloud, but execute queries locally. | Eliminates the retrieval latency tax by removing network round trips. |
| Memory | Keep session memory local. Load long term memory once at session start and synchronize asynchronously. | Avoids paying another network hop on every conversational turn. |
| Sessions | Isolate each conversation with a single authoritative source of state. Snapshot sessions for recovery and reconnects. | Prevents state leakage while making conversations resilient. |
| Orchestration | Keep the execution loop simple. Everything on the hot path should be a local function call whenever possible. | Latency compounds across tool calls. |
| Deployment | Treat the cloud as the control plane and the browser, edge, or device as the data plane. | Execute latency sensitive work as close to the user as possible. |

The architecture follows a few simple principles:

- Route models instead of standardizing on one.
- Treat retrieval as a local operation, not a network service.
- Keep conversational state inside the runtime.
- Move network calls off the critical path whenever possible.
- Separate the control plane from the data plane.
- Optimize for time to first token, not benchmark scores.

The end result isn't a fundamentally different AI stack. It's the same seven layers, reorganized around the constraint that matters most in production: latency. Once retrieval, memory, and session state become local operations instead of distributed systems problems, the entire architecture becomes simpler, faster, and more predictable.

## The Future of the Production AI Stack

The last two years have been dominated by the race for better models. Every release has improved reasoning, expanded context windows, lowered inference costs, or pushed benchmark scores higher, leading many teams to assume that choosing the right model is the primary architectural decision when building an AI product. In practice, however, production systems are increasingly limited by everything surrounding the model rather than the model itself.

Once an application reaches production, latency, reliability, and consistency are determined less by which LLM generates the response and more by how quickly the system can retrieve knowledge, access memory, coordinate tools, maintain session state, and move information between components. Every network hop added to that execution path compounds across a conversation, eventually becoming the dominant factor in the user experience. Improving the model cannot recover time that has already been spent waiting on infrastructure.

That is why we believe the next generation of AI infrastructure will look fundamentally different from the web architectures it inherited. Instead of treating retrieval, memory, and state as remote services accessed over the network on every request, production AI systems will increasingly execute those operations where the conversation is actually happening, whether that's in the browser, on the edge, on device, or alongside the application itself. The cloud doesn't disappear, but its role shifts toward building indexes, synchronizing state, coordinating deployments, and acting as the system of record, while the latency sensitive execution path moves as close to the user as possible.

This architectural shift is what led us to build Moss.

Rather than treating search as another hosted service, Moss executes hybrid semantic and keyword retrieval directly inside the runtime, allowing agents to retrieve knowledge in single digit milliseconds without paying the retrieval latency tax that has become standard across much of the industry. The goal isn't simply to build a faster search engine, but to make retrieval behave like any other local function call so that developers can ground every meaningful interaction without compromising responsiveness.

If the central argument of this article is that production AI systems should be designed around latency budgets rather than infrastructure defaults, then Moss is our implementation of that philosophy. We believe the fastest AI systems won't be the ones with the best models alone, but the ones whose surrounding infrastructure is architected so efficiently that the model can begin reasoning almost immediately. That's the stack we're building toward, and we think it's where production AI is headed.
]]></content:encoded>
            <author>Sri Raghu Malireddi</author>
            <author>Ashvath Suresh Kumar</author>
            <category>production AI architecture</category>
            <category>AI runtime</category>
            <category>conversational AI architecture</category>
            <category>reference architecture</category>
            <category>AI infrastructure</category>
            <category>retrieval latency</category>
            <category>RAG</category>
            <category>voice AI</category>
            <category>edge AI</category>
            <category>AI agents</category>
            <enclosure url="https://mossdev.work/blog/the-production-ai-stack/og.png" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[We Built a Voice AI Agent for Our Website. Then Other Companies Started Asking for It.]]></title>
            <link>https://mossdev.work/blog/founding-agent</link>
            <guid isPermaLink="false">https://mossdev.work/blog/founding-agent</guid>
            <pubDate>Fri, 12 Jun 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[We built a voice AI agent for mossdev.work so visitors could get answers without waiting for a sales call. It quickly became one of our most effective ways to engage leads, and other companies started asking for it. Meet Founding Agent.]]></description>
            <content:encoded><![CDATA[
As founders, there are a handful of questions we answer over and over again.

What does your company do? How is it different? How does the technology work? Can it integrate with my stack? How much does it cost? Do you support enterprise?

The interesting thing is that most of these conversations never happen on a sales call. They happen before the sales call.

Someone lands on your website. They're interested enough to learn more, but not interested enough to book a meeting. So they click around. They skim product pages, documentation, FAQs, and blog posts looking for answers.

Sometimes they find what they need. Sometimes they don't.

And when they don't, they leave.

For years, we've accepted this as the normal website experience. We started wondering why.

Why is a website still mostly a collection of pages? Why can't it simply answer questions? Why can't visitors ask what they want to know and get an immediate response?

That question led us to build something for ourselves.

## Built to Solve Our Own Problem

At Moss, we're obsessed with retrieval speed. It's the foundation of everything we do.

Most retrieval systems were designed around search. We wanted retrieval that felt conversational, fast enough that someone could ask a question and get an answer instantly.

So we built a voice AI agent and put it directly on our website.

Visitors could simply start talking. No forms. No chat windows. No digging through documentation. Just a conversation.

Initially, it was an experiment. We wanted to see if we could make the experience feel natural and helpful. Could an AI agent answer questions in real time and feel less like software and more like talking to someone on the team?

The answer was yes.

## The Unexpected Result

Once the agent was live, visitors started asking the same questions we answer every day about our infrastructure, integrations, pricing, implementation, and product capabilities.

The difference was that we weren't answering them anymore.

The agent was.

Because it was grounded in our website, documentation, FAQs, presentations, and internal knowledge, it didn't sound like a generic chatbot. It sounded like Moss.

The result was simple: less repetitive work, better qualified conversations, and more time spent talking to people who were genuinely ready to move forward.

## Then People Started Asking for It

This is often how products are born.

Not because you set out to build one, but because people keep asking if they can have what you built for yourself.

Visitors would interact with the agent and then ask:

"Can I put this on my website?"

"How difficult was this to build?"

"Can it answer questions about my company?"

"Can it qualify prospects for my team?"

The more conversations we had, the more obvious it became that this wasn't just a Moss problem.

Every company has visitors with questions. Every founder answers the same things repeatedly. Every team has prospects who are interested but not quite ready to book a meeting.

Yet most websites still ask visitors to fill out a form and wait for a response.

That feels like a broken experience.

## Introducing Founding Agent

Founding Agent is a voice AI agent that lives on your website.

It answers visitor questions, qualifies prospects, and gives every visitor an immediate conversation instead of a static browsing experience.

It's trained on your company, products, documentation, FAQs, and expertise. When someone asks a question, they get the answer your team would have given. Not a generic AI response. Not a hallucinated one.

And because it's powered by Moss's retrieval infrastructure, responses happen in milliseconds, making conversations feel natural and responsive.

The way they should.

## Try It Yourself

The best way to understand Founding Agent isn't to read about it.

It's to talk to it.

We've put it directly on our website, so ask it anything. Pricing questions. Technical questions. Product questions. Difficult questions.

See how it responds.

If you're interested in bringing Founding Agent to your own website, we're opening early access and looking for teams that want to give every visitor an answer the moment they ask for one.

Instead of making them wait.

<div className="flex justify-center">
  <a
    href="https://mossdev.work/fa"
    className="inline-flex h-[46px] items-center rounded-full bg-moss-200 px-5 text-sm leading-[1.3] text-fg-default no-underline transition-[transform,background-color] duration-200 hover:-translate-y-px hover:bg-moss-300"
  >
    <span>Join the waitlist →</span>
  </a>
</div>
]]></content:encoded>
            <author>Sri Raghu Malireddi</author>
            <category>Founding Agent</category>
            <category>voice AI</category>
            <category>conversational AI</category>
            <category>lead qualification</category>
            <enclosure url="https://mossdev.work/blog/founding-agent/og.png" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[What Happens When You Remove the Network Hop from RAG]]></title>
            <link>https://mossdev.work/blog/remove-the-network-hop</link>
            <guid isPermaLink="false">https://mossdev.work/blog/remove-the-network-hop</guid>
            <pubDate>Tue, 17 Mar 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[We profiled a production RAG pipeline before and after eliminating the retrieval network round-trip. The results explain why your AI agent feels slower than it should.]]></description>
            <content:encoded><![CDATA[
In our [last post](/blog/retrieval-latency-tax), we made the case that retrieval, not the LLM, is the dominant latency bottleneck in real-time AI applications. We showed the numbers. We named the problem.

This post is about what happens when you fix it.

We took a production RAG pipeline, a voice agent doing knowledge-base lookups over a managed vector database, and ran a controlled experiment. Same data. Same queries. Same embedding model. Same LLM. The only variable: where the retrieval happens. In one configuration, retrieval goes over the network to a cloud-hosted vector database. In the other, the index lives in the same process as the agent, and retrieval is a local function call.

The difference isn't incremental. It's architectural.

## The Baseline: A Typical Cloud RAG Pipeline

Here's the setup we profiled. It's not a strawman. It's what most production RAG applications look like today.

A voice agent running on a cloud VM receives transcribed speech from a streaming ASR service. It sends an embedding request to an embedding API, receives the vector back, queries a managed vector database (in this case, hosted in the same cloud region), gets the top-k results, assembles the prompt with retrieved context, and sends it to an LLM for generation. The generated text streams to a TTS service for audio synthesis.

We instrumented every hop. Here's the median latency breakdown for the retrieval portion alone, from the moment the agent has the user's transcribed text to the moment it has retrieved context ready for prompt assembly:

| Step | Median | P95 | P99 |
|---|---|---|---|
| Embedding API call | 22ms | 38ms | 67ms |
| Network to vector DB | 12ms | 24ms | 51ms |
| Vector search (DB processing) | 18ms | 31ms | 44ms |
| Network return | 11ms | 22ms | 48ms |
| Deserialization + re-ranking | 4ms | 7ms | 12ms |
| **Total retrieval** | **67ms** | **122ms** | **222ms** |

67ms median looks manageable. But look at the P99: **222ms for a single retrieval call.** And remember, sophisticated agents make two to three retrieval calls per turn. At P99, that's **450–670ms** just in retrieval, before the LLM has seen a single token.

This is the tail latency trap. Your median looks fine. Your P99 is destroying your user experience. And in voice applications, [tail latency *is* the experience](https://www.rack2cloud.com/deterministic-networking-ai-infrastructure/), because users don't perceive averages. They perceive the worst moments.

## What Changes When You Go Local

Now, the same pipeline with retrieval co-located in the agent process. The index, a compact, pre-built vector index, is loaded into the agent's memory at startup. When the agent needs to retrieve, it calls a local function. No network. No serialization. No connection pools.

| Step | Median | P95 | P99 |
|---|---|---|---|
| Embedding (local model) | 3ms | 5ms | 8ms |
| Vector search (in-process) | 1.2ms | 2.1ms | 3.4ms |
| Re-ranking | 0.8ms | 1.4ms | 2.1ms |
| **Total retrieval** | **5ms** | **8.5ms** | **13.5ms** |

Read that again. Median retrieval dropped from **67ms to 5ms**. P99 dropped from **222ms to 13.5ms**. That's a **13x improvement at median** and a **16x improvement at P99**.

> **Three retrieval calls per turn: 15ms total instead of 670ms at P99. Over half a second reclaimed.**

That's enough to transform a sluggish voice agent into one that feels instantaneous.

## Where the Time Actually Disappeared

The numbers are dramatic, but the *why* matters more than the *what*. Let's trace where the latency evaporated.

**The network round-trip vanished.** This is the obvious one, but it's worth quantifying. In the cloud baseline, the network contributed 23ms at median and 99ms at P99, just for the two hops to the vector database and back. In the same region. On a fast network. With keep-alives and connection pooling. The local configuration has zero network latency because there is no network. The data is in the same address space as the code that needs it.

**Serialization overhead disappeared.** Every network call requires serializing the query (embedding vector + filter parameters) into a wire format, transmitting it, deserializing on the database side, processing, serializing the results, transmitting back, and deserializing again. With in-process retrieval, the query is a function call with a pointer to the vector. The results are a pointer to the matches. No copying. No encoding. No protocol overhead.

**Connection management evaporated.** Managed vector databases require connection pools, authentication tokens, TLS handshakes, and retry logic. These are well-engineered systems, but they add overhead, especially at P99, where you occasionally hit a cold connection, a pool exhaustion event, or a TLS renegotiation. Local retrieval has none of this. The index is a data structure in memory. You call a function. It returns.

**The embedding step shrank.** In the cloud baseline, embedding required an API call to an external service: 22ms median, 67ms at P99. With a co-located lightweight embedding model (quantized, optimized for the target hardware), the same embedding operation takes 3ms. The model is smaller, yes, but for retrieval queries (short text, not documents), a well-optimized compact model achieves nearly identical recall at a fraction of the latency.

**Tail latency collapsed.** This is the most underappreciated benefit. Cloud services have inherently variable latency: network jitter, garbage collection pauses on the database, load balancer rebalancing, noisy neighbors on shared infrastructure. These factors don't affect median much, but they blow up P99. In-process retrieval on dedicated hardware has near-deterministic latency. The P99/P50 ratio dropped from 3.3x (cloud) to 2.7x (local). A much tighter distribution.

## The Compounding Effect

Here's what happens to the full voice agent pipeline when retrieval drops from 200ms+ to under 15ms:

| Component | Cloud RAG (P95) | Local RAG (P95) |
|---|---|---|
| ASR | 150ms | 150ms |
| Retrieval (x2 calls) | 244ms | 17ms |
| LLM (TTFT) | 280ms | 280ms |
| TTS (first audio) | 120ms | 120ms |
| **Total to first audio** | **794ms** | **567ms** |

That 227ms difference is the gap between an agent that *barely* meets the [800ms voice interaction benchmark](https://www.pnas.org/doi/10.1073/pnas.0903616106) and one that sails under it with room to spare. Room for an extra retrieval call. Room for a more capable (slower) LLM. Room for re-ranking, safety checks, or citation generation.

> **Fast retrieval doesn't just make retrieval better. It gives you back architectural headroom to invest in everything else.**

## "But Will It Scale?"

The immediate objection is obvious: an in-process index can't hold as much data as a cloud database. This is true, and it's the wrong frame.

The question isn't whether a local index can replace a cloud database for every workload. It's whether the *retrieval that happens on the hot path*, the queries that sit between user input and agent response, needs to go over a network.

Most voice agents and copilots retrieve from knowledge bases that range from tens of thousands to a few million chunks. A well-compressed vector index for 1 million 256-dimensional vectors occupies roughly 500MB. That's well within the memory budget of any modern server, and feasible even for browser-based WASM runtimes with aggressive compression.

The pattern that works: **keep the network database as the system of record, but distribute compact, pre-built indexes to the runtimes that need them.** The index at the edge is a read-optimized projection of the data, not a replacement for the primary store. Updates flow from the source through an indexing pipeline, and fresh indexes are distributed to runtimes on a cadence that matches the data's rate of change.

This is the same pattern the web itself runs on. CDNs don't replace origin servers. They distribute read-optimized copies to the edge so that the hot path (serving a page to a user) doesn't pay the round-trip to the origin. The insight is identical:

> **Move the data closer to where it's consumed, not the consumer closer to the data.**

## The Queries That Don't Need a Network

Not every query should go local. Analytical queries over your full corpus ("find all documents tagged 'compliance' from the last quarter") belong in a server-side database. Complex joins across multiple indexes, aggregations, and batch processing are server workloads.

But the queries that AI agents make in real time follow a predictable pattern:

**Short query vectors.** User utterances, typically 5–30 tokens, embedded into a single vector. The query payload is tiny.

**Small result sets.** Top-5 or top-10 chunks. The agent doesn't need the full corpus ranked. It needs a handful of highly relevant passages to inject into the prompt.

**Repeated over narrow scopes.** A voice agent handling customer support queries for a specific product retrieves from a knowledge base of maybe 10,000–50,000 chunks. A copilot autocompleting code retrieves from a repo of maybe 100,000 chunks. These are small, bounded corpora.

**Latency-critical.** Every millisecond between the user's input and the agent's response is perceived. These queries are the definition of hot-path operations.

This profile is a perfect match for in-process retrieval. Small index, small queries, small results, extreme latency sensitivity. Sending these queries over a network is like routing every function call through a REST API: architecturally possible, but the overhead dominates the actual work.

## What Changes in the Developer Experience

Beyond raw latency, removing the network hop simplifies the entire developer workflow.

**No infrastructure to manage.** No database cluster to provision, scale, monitor, or pay for by the hour. The index is a file. You load it at startup. If the process restarts, it reloads.

**No connection strings.** No configuring regions, authentication, connection pools, retry policies, or timeout values. No debugging why retrieval is slow at 3 AM because the connection pool is exhausted.

**Deterministic testing.** Your CI pipeline can load the same index file and run the same queries with the same results. No flaky tests because the hosted database had a latency spike. No mocking retrieval calls for unit tests. Just call the real thing.

**Offline capability.** If your agent runs on a device or in a browser, local retrieval works without an internet connection. The knowledge base traveled with the runtime. This isn't a niche requirement. It's table stakes for mobile applications, aircraft systems, field service tools, and any deployment where connectivity is intermittent.

## The Shift in Mental Model

The deeper change isn't technical. It's conceptual. For the past several years, the default mental model for retrieval in AI applications has been: *"Retrieval is a network service you call."* This model was inherited from how we've always built database-backed applications. It was never questioned because it's how databases work.

But retrieval in an AI agent isn't the same as a database query in a web application. A web app makes one or two database calls per page load, with a latency budget of 200ms that users won't notice. An AI agent makes multiple retrieval calls per conversational turn, with a latency budget of 50ms or less if it wants to feel responsive.

> **The new mental model: retrieval is a function you call, not a service you query.**

Same semantics. Same results. Fundamentally different performance characteristics. The index is a local data structure, not a remote service. Querying it is a function call, not a network request.

This is the same conceptual shift that happened when [SQLite](https://sqlite.org/whentouse.html) proved that not every application needs a client-server database. Not every retrieval needs a client-server vector store. The workload characteristics of real-time AI retrieval (small, fast, repeated, latency-critical) are precisely the characteristics that favor embedded, co-located data structures over network services.

## What This Means for What You're Building

If you're building a real-time AI application (a voice agent, a copilot, a conversational search product), run this experiment yourself.

**Profile your retrieval path.** Measure the full round-trip, including embedding, network transit, search, and deserialization. Compare median to P99. If P99 is more than 3x your median, network variability is dominating your tail latency.

**Calculate your retrieval budget.** Take your total acceptable turn latency (800ms for voice, maybe 1.5s for text), subtract ASR, LLM, and TTS. Whatever's left is your retrieval budget. If your current retrieval exceeds it, the architecture is the bottleneck, not the implementation.

**Estimate your index size.** Count your chunks, multiply by your embedding dimensions times 4 bytes (for float32) or 1 byte (for int8 quantized). If the result fits in memory (and for most agent workloads, it will), the data can live in-process.

The network hop in your RAG pipeline isn't a fixed cost. It's an architectural choice. And for the class of workloads that defines the next generation of AI applications (real-time, conversational, latency-critical), it's a choice worth reconsidering.

> **The retrieval doesn't need to be faster. It needs to be *closer*.**

---

*Sri Raghu Malireddi is the Founder & CEO of Moss. Previously ML Lead at Grammarly and Microsoft, where he built real-time ML systems serving millions of users. His research has been published at ACL and NAACL, and he is an inventor on multiple patents in real-time machine learning. [LinkedIn](https://linkedin.com/in/r4ghu) · [X](https://x.com/srimalireddi)*
]]></content:encoded>
            <author>Sri Raghu Malireddi</author>
            <category>RAG optimization</category>
            <category>retrieval latency</category>
            <category>embedded search</category>
            <category>AI infrastructure</category>
            <category>voice AI</category>
            <enclosure url="https://mossdev.work/blog/remove-the-network-hop/og.png" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[The Retrieval Latency Tax: Why Your AI Agent Feels Slow (And It's Not the LLM)]]></title>
            <link>https://mossdev.work/blog/retrieval-latency-tax</link>
            <guid isPermaLink="false">https://mossdev.work/blog/retrieval-latency-tax</guid>
            <pubDate>Tue, 17 Mar 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Everyone blames the model. But in real-time AI, the real bottleneck is retrieval. Here's the data that proves it, and what it means for the next generation of AI applications.]]></description>
            <content:encoded><![CDATA[
Ask any developer why their AI agent feels slow, and you'll get the same answer: *"The model is too slow."*

It's intuitive. Language models are computationally expensive. Inference takes time. Tokens stream in one by one. So when an AI assistant takes a second to respond, or worse, when a voice agent goes silent for a beat too long, the LLM gets the blame.

But it's wrong. Or at least, it's incomplete.

We've spent the past year profiling real-time AI applications: voice agents, retrieval-augmented copilots, conversational search systems. The data tells a different story. In most architectures, **retrieval is the dominant source of user-facing latency**, not the language model. And unlike LLM inference, which is improving rapidly with every generation of hardware and model optimization, retrieval latency has barely moved in three years.

This is the retrieval latency tax. Every AI agent pays it. Almost nobody talks about it. And it's quietly killing the user experience of the most promising AI applications being built today.

## The Anatomy of an AI Agent Turn

To understand where the time goes, you need to decompose a single turn of an AI agent, the cycle from user input to agent response. Here's what a typical RAG-powered voice agent looks like:

**Step 1: Speech Recognition (ASR).** The user speaks. An automatic speech recognition model transcribes the audio to text. Modern streaming ASR from providers like Deepgram or AssemblyAI adds roughly **100–200ms** of latency, depending on utterance length and endpoint detection.

**Step 2: Retrieval.** The transcription triggers a retrieval call. The agent needs context: a knowledge base article, a product spec, a conversation history snippet. This query gets embedded, sent to a vector database over the network, the database performs similarity search, and the results travel back. Best case: **250–500ms**. Realistic with network variability, cold starts, and multiple retrievals per turn: **500ms–1.5s**.

**Step 3: LLM Inference.** The retrieved context, system prompt, and user query are assembled and sent to the language model. With streaming, the first token typically arrives in **200–400ms**, and the full response generates over the next few hundred milliseconds.

**Step 4: Speech Synthesis (TTS).** The generated text is converted to audio. Modern streaming TTS from ElevenLabs, Cartesia, or PlayHT begins playback in **75–200ms**.

Add it up: **800ms to 1.5 seconds** before the agent can even *begin* speaking. And that's generous. That assumes a single retrieval call, no cache misses, and stable network conditions.

## The Number That Matters: 300 Milliseconds

[Decades of conversational analysis research](https://www.pnas.org/doi/10.1073/pnas.0903616106) have established that humans perceive pauses longer than roughly **300 milliseconds** as unnatural in dialogue. Beyond 500ms, the pause registers as the other party being confused, disengaged, or struggling to respond. Beyond a second, most people start to disengage entirely.

This isn't a preference. It's deeply wired into how human conversation works. When your voice agent takes 800ms to a full second before uttering its first syllable, the user isn't thinking *"the model is processing my request."* They're thinking *"this thing is broken."* Or they've already hung up.

Voice AI platforms know this. Retell AI has reported average response times of approximately **800ms** across their platform. Synthflow has documented latencies as low as **420ms** in optimized conditions. The industry consensus is converging around **800ms as the benchmark** for acceptable voice agent response time, and most implementations struggle to hit it consistently.

## Where the Time Actually Goes

Here's what makes the retrieval latency tax so hard to fix: it's invisible in standard benchmarks.

When developers optimize their AI agents, they focus on what's measurable and attributable. LLM latency is highly visible. Every inference API returns timing headers. Model providers compete on time-to-first-token. Teams benchmark GPT-4 Turbo against Claude against Gemini, comparing milliseconds.

But retrieval latency hides in the gaps. It's spread across:

**Network round-trips.** Your agent runs in us-east-1. Your vector database runs in us-west-2. Or your agent runs in a user's browser, and the vector database runs... anywhere else. Every query pays the network tax twice, once out, once back. Even within the same cloud region, you're looking at 10–50ms of network overhead per call. Across regions or from edge to cloud, it's 50–200ms.

**Cold starts and connection overhead.** Managed vector databases have connection pools, authentication handshakes, and occasional cold starts. If your agent hasn't queried in a while, that first retrieval can be significantly slower than the steady-state.

**Query processing.** The database itself takes time. Embedding the query (if it doesn't arrive pre-embedded), performing approximate nearest neighbor search, filtering results, re-ranking, and serializing the response. Published p99 latencies from major vector databases tell the story: Pinecone reports around **45ms p99**, Qdrant approximately **35ms p99**, Weaviate roughly **65ms p99**. These are good numbers for the database. But they don't include the network round-trip that wraps every query.

**The multiplication problem.** This is where it gets really painful. Sophisticated AI agents don't make one retrieval call per turn. They make **two to three**. A voice agent might first retrieve from a knowledge base, then pull conversation history, then check a policy document. Each retrieval call pays the full tax: network, processing, and return. Three calls at 150–300ms each puts retrieval at **450ms–900ms per turn**. That's before the LLM has seen a single token.

## The Benchmarks Everyone Ignores

The AI industry has gotten exceptionally good at benchmarking models. We have MMLU, HumanEval, MT-Bench, Chatbot Arena. Dozens of standardized ways to compare language models on quality and speed.

We have almost nothing equivalent for retrieval latency in agent workflows.

This is a massive blind spot. Teams will spend weeks evaluating whether to use GPT-4o or Claude Sonnet for a 50ms difference in time-to-first-token, while ignoring the 300–500ms of retrieval latency sitting in the same pipeline. They'll optimize prompt length to shave 100ms off inference, then send three round-trip network calls to a database that adds 600ms.

> **Retrieval is the highest-leverage latency target in most AI agent architectures today.**

[ElevenLabs proved this](https://elevenlabs.io/blog/engineering-rag) when they optimized their conversational AI pipeline. By restructuring their RAG implementation, they reduced retrieval latency from **326ms to 155ms**, a 52% improvement. The result wasn't incremental. It fundamentally changed the feel of their voice agents. Not because the model got smarter. Because the plumbing got faster.

## Why This Problem Is Getting Worse, Not Better

Three trends are compounding the retrieval latency tax:

**Agents are getting more autonomous.** The era of single-turn Q&A is ending. Modern AI agents execute multi-step workflows: researching, planning, acting, and iterating. Each step often requires fresh context retrieval. An agent that makes 5 retrieval calls per workflow at 200ms each adds a full second of latency before any model inference happens.

**Voice is becoming the primary interface.** Text-based chatbots can hide latency behind typing indicators and streaming text. Voice agents can't. Dead air is dead air. As conversational AI shifts toward voice-first interactions (customer service, healthcare, sales, accessibility), the tolerance for retrieval latency drops to near zero.

**Edge and browser deployments are growing.** AI is moving out of centralized cloud servers and into browsers, mobile apps, and edge devices. This is great for privacy and user experience, but it makes the retrieval problem worse by an order of magnitude. If your vector database lives in the cloud and your agent runs in a user's browser, every retrieval call now pays the full internet round-trip penalty. There's no same-region optimization to fall back on.

## The Architecture Problem

> **The retrieval latency tax isn't a bug. It's a fundamental property of the architecture.**

Every major vector database (Pinecone, Qdrant, Weaviate, Milvus, Chroma) is designed as a *network service*. You deploy it, or someone deploys it for you, and your application communicates with it over HTTP or gRPC. This architecture was inherited from traditional databases, and it makes perfect sense for traditional database workloads.

But AI agent retrieval isn't a traditional database workload. It's not a batch analytics query. It's not a server-side API call where 200ms is invisible. It's a hot-path operation that sits directly between a user's input and an AI's response, repeated multiple times per turn, where every millisecond of delay is perceived as reduced intelligence.

We wrote about this principle in our [first post](https://mossdev.work/blog/we-spent-a-decade-making-ai-feel-instant):

> **In any interactive AI system, perceived intelligence is bounded by perceived speed.**

You can optimize the database query all you want. You can compress embeddings, use quantization, add caching layers, pre-fetch likely queries. These are all worthwhile optimizations. But they cannot eliminate the network hop. And as long as retrieval means "send a request over a network and wait for a response," there's a floor to how fast it can get.

The database providers know this. That's why they've invested heavily in lower-latency networking, edge deployments, and connection pooling. These are real improvements. But they're optimizing within the constraints of a network-service architecture: making the round-trip faster, not eliminating it.

## What the Next Architecture Looks Like

The solution to the retrieval latency tax isn't a faster database. It's **moving retrieval out of the network path entirely.**

What if the index lived in the same process as the agent? No network hop. No connection overhead. No cold starts. No multiplication penalty, because a local lookup takes microseconds, not milliseconds, regardless of how many you make per turn.

This isn't a theoretical idea. It's the direction the industry is heading. The same way that SQLite proved you don't always need a client-server database, the AI agent ecosystem is discovering that you don't always need a client-server vector store. For real-time workloads (voice agents, copilots, conversational search), the retrieval layer should be *co-located* with the inference layer. Not nearby. Not in the same region. In the same process.

The engineering challenges are real: you need compact index formats that fit in constrained environments, efficient vector search that runs in single-digit milliseconds without dedicated hardware, and a sync mechanism that keeps distributed indexes fresh. But these are solvable problems, and solving them removes the single biggest source of user-facing latency in modern AI applications.

## Measuring Your Retrieval Tax

If you're building a real-time AI application, here's a quick diagnostic:

**Instrument your retrieval calls.** Not just the database query time. Measure the full round-trip from the moment your agent code initiates the retrieval to the moment it has results in memory. Include connection acquisition, serialization, network transit, and deserialization.

**Count your retrievals per turn.** Most agents make more retrieval calls than developers realize, especially if you're using frameworks like LangChain or LlamaIndex that abstract retrieval behind chains and tools.

**Calculate your retrieval tax as a percentage.** Take your total retrieval time per turn and divide it by your total turn latency. If retrieval accounts for more than 30% of your end-to-end latency, it's your highest-leverage optimization target.

**Test from the user's location.** Retrieval latency benchmarks from within the same data center are meaningless if your users are on the other side of the continent, or if your agent runs on their device.

## The Conversation We Need to Have

The AI industry is in the middle of a massive investment in model intelligence. Billions of dollars flowing into foundation models, reasoning capabilities, multimodal understanding, and agent frameworks. This is important work.

> **But intelligence without speed is a product that nobody uses.**

The best AI agent in the world, the one with perfect retrieval quality, flawless reasoning, and empathetic responses, will lose to a mediocre agent that responds in 400ms instead of 1,200ms. Not because users are impatient (though they are). Because perceived speed *is* perceived intelligence. A fast response feels smart. A slow response feels broken.

The retrieval latency tax is the largest unsolved performance problem in real-time AI. It's not glamorous. It doesn't make for exciting model announcements or benchmark leaderboard victories. But for the teams actually building voice agents, copilots, and conversational AI products, the ones where user experience is measured in milliseconds, it's the problem that determines whether their product feels magical or frustrating.

> **The models are fast enough. The question is whether the plumbing is.**

---

*Sri Raghu Malireddi is the Founder & CEO of Moss. Previously ML Lead at Grammarly and Microsoft, where he built real-time ML systems serving millions of users. His research has been published at ACL and NAACL, and he is an inventor on multiple patents in real-time machine learning. [LinkedIn](https://linkedin.com/in/r4ghu) · [X](https://x.com/srimalireddi)*
]]></content:encoded>
            <author>Sri Raghu Malireddi</author>
            <category>retrieval latency</category>
            <category>voice AI</category>
            <category>RAG performance</category>
            <category>conversational AI</category>
            <category>AI infrastructure</category>
            <enclosure url="https://mossdev.work/blog/retrieval-latency-tax/og.png" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[We Spent a Decade Making AI Feel Instant. Here's What We Learned.]]></title>
            <link>https://mossdev.work/blog/we-spent-a-decade-making-ai-feel-instant</link>
            <guid isPermaLink="false">https://mossdev.work/blog/we-spent-a-decade-making-ai-feel-instant</guid>
            <pubDate>Tue, 10 Mar 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Moss founder and CEO Sri Raghu Malireddi shares why he started Moss, and why the future of real-time AI depends on retrieval becoming a runtime instead of another network service.]]></description>
            <content:encoded><![CDATA[
Last fall, I was prototyping an AI agent. The RAG pipeline was solid: good embeddings, well-indexed corpus, decent retrieval quality. On paper, it worked. In practice, every user interaction followed the same pattern: the agent needed context, so it called out to a vector database over the network. That call took anywhere from 300ms to a couple of seconds depending on load and location. The agent was making two to three retrieval calls per turn. By the time it had the context it needed, the user had been waiting over a second before anything useful started happening.

In a chatbot, a second is annoying. In a voice agent, it's a dead pause that makes people hang up. In a copilot, it's long enough for the user to context-switch to another tab and lose their train of thought.

I'd spent years watching this exact problem kill otherwise good products. This post is about why I started Moss and why I became convinced retrieval needed a completely different architecture.

## Speed as the Ceiling on Intelligence

Before Moss, I was an ML Lead at Grammarly, working on real-time writing assistance for 40 million daily active users. Before that, I built ML systems for Bing and Office at Microsoft.

At Grammarly, I led personalization for Grammarly Keyboard: making AI suggestions feel right on a mobile device where every millisecond counts. We ran models on-device, ranked suggestions in real-time, and optimized until the keyboard felt like it was reading your mind. That work drove 300% retention growth. The model was already good. What changed was that users could actually feel how good it was, because the suggestions arrived before they lost patience.

A principle came out of that: **in any interactive AI system, perceived intelligence is bounded by perceived speed.** I published research at ACL and NAACL, filed multiple patents in real-time ML. But the most useful thing I learned was watching real users and seeing exactly when latency killed the magic.

We solved this for writing suggestions at Grammarly. But conversational AI (voice agents, copilots, anything with back-and-forth rhythm) hit a wall that model optimization alone couldn't fix.

## The Problem

During early prototypes, one thing became obvious. The retrieval quality wasn't the bottleneck anymore. Latency was.

Every conversational turn required multiple network round trips to retrieve context. The embeddings were good. The ranking was good. The models were good. But the user still waited.

That's when I realized the bottleneck wasn't semantic search itself. It was the assumption that retrieval always had to happen over the network.

Modern AI infrastructure had accepted that assumption as inevitable. I wasn't convinced it was.

It wasn't just a performance problem. Developers building AI agents were cobbling together vector databases, sync pipelines, caching layers, and embedding services, then spending more time on plumbing than on the product. What should be a single operation (*find the relevant context and return it*) had become a multi-service architecture problem.

Vector databases like Pinecone, Weaviate, and Qdrant are good at what they do. But they were designed for a world where the query originates from a server, crosses a network, hits a managed cluster, and returns. That works for offline analytics or server-side RAG. It breaks down for voice agents making multiple lookups per turn, copilots that need instant recall, or anything where the user is waiting.

The problem isn't that these databases are slow. The problem is a network hop baked into the architecture. No amount of optimization on the database side eliminates that.

## The Question That Changed Everything

Back to that voice agent prototype. In the early days of Moss, we tried caching, connection pooling, pre-fetching, embedding compression. Shaved off 50ms here, 30ms there. But as long as retrieval meant "send a request over a network and wait," we were fighting physics.

Then a simple question reframed everything: *What if retrieval didn't happen over the network at all? What if the index lived in the same process as the agent?*

Taking that question seriously means retrieval stops being a service you query and becomes a runtime: indexing, synchronization, and local semantic search that ship with the agent and run wherever the agent runs, whether that's a server, the edge, a browser, or a device.

Making that work requires an index format compact enough to distribute to browsers and edge devices, a runtime fast enough for vector similarity search in single-digit milliseconds in WebAssembly, a sync layer that keeps distributed indexes fresh without rebuilding them, and all of it packaged as a single `pip install` or `npm install`.

That's what Moss set out to build.

## Why Rust and WebAssembly

If the search runtime has to live inside the agent process (Node.js server, Python backend, browser tab, mobile app) you need something that compiles to every target, runs at near-native speed, and has a small memory footprint.

C++ has the performance but painful WebAssembly tooling, memory safety liabilities, and rough developer experience. JavaScript gives portability but not performance: vector math in JS is an order of magnitude slower than native.

Rust gave us native speed, memory safety without garbage collection, first-class WebAssembly compilation via `wasm-pack`, and a type system that catches entire categories of bugs at compile time. The Rust core is the single source of truth. From it, we cross-compile to WebAssembly for browsers and edge runtimes, and generate native Python and TypeScript bindings. Developers get an SDK that feels native to their language, but under the hood it's the same Rust engine everywhere.

The result: one codebase that runs identically in a Python process, a Node.js server, a browser tab, a Cloudflare Worker, or a React Native app. Same code, same performance, same API.

## What Moss Does

Moss is a real-time retrieval runtime. We coined that term because nothing existing described what we were making. Not a database (we don't store your data). Not a RAG framework (we don't orchestrate your LLM pipeline). We're the layer that makes retrieval instant and local, wherever your agent runs.

You connect your data source once. Moss indexes it, creates a compact distributable artifact, and pushes it to wherever your agent lives. When your agent needs context, it does a local lookup in under 10 milliseconds. No network hop. No cold start. No external dependency.

`pip install moss`, point it at your data, and retrieval just works.

## What's Next

Since then, Moss has grown from an idea into production infrastructure powering tens of millions of real-time AI interactions every month. Developers use Moss across cloud, edge, browser, and on-device environments where every millisecond matters.

But we're still at the beginning. Retrieval is becoming one of the fundamental primitives of real-time AI systems, and we're building the runtime that makes it effectively disappear.

This post is the first in a series. Coming up:

- **The Retrieval Latency Tax**: where the bottlenecks actually live in AI agent architectures
- **Inside Moss's Architecture**: how we built sub-10ms semantic search in Rust and WebAssembly
- **Benchmarks**: reproducible performance comparisons against cloud vector databases on real agent workloads

If retrieval latency is on your critical path, [try Moss](https://portal.usemoss.dev) or come talk to us on [Discord](https://moss.link/discord).

---

*Sri Raghu Malireddi is the Founder & CEO of Moss. Previously ML Lead at Grammarly and Microsoft, where he built real-time ML systems serving millions of users. His research has been published at ACL and NAACL, and he is an inventor on multiple patents in real-time machine learning. [LinkedIn](https://linkedin.com/in/r4ghu) · [X](https://x.com/srimalireddi)*
]]></content:encoded>
            <author>Sri Raghu Malireddi</author>
            <category>founding story</category>
            <category>real-time AI</category>
            <category>conversational AI</category>
            <category>voice AI</category>
            <enclosure url="https://mossdev.work/blog/we-spent-a-decade/og.png" length="0" type="image/png"/>
        </item>
    </channel>
</rss>