Aside: AI memory on 80K+ devicesHow Aside powers AI memory across 80,000+ devices · 113M documents · 24.8B tokens · 150+ countries

Read moreRead the case study
Moss
usemossStart Free
  1. Integrations
  2. /
  3. Pipecat

Moss + Pipecat

Real-Time Search for Voice Pipelines

Get StartedTalk to an Engineer

About Pipecat

Pipecat is an open-source framework for building voice and multimodal agents. The pipecat-moss package provides MossRetrievalService, a pipeline processor that queries Moss on every user turn and injects context before the LLM generates a response. No tool calling needed. Sub-10ms retrieval keeps conversations flowing without dead air.

Why Use Moss with Pipecat

01

Native Pipecat pipeline processor

MossRetrievalService slots between STT and LLM in your pipeline

02

Context injection, no tool-calling overhead

Search runs automatically on every user utterance

03

Sub-10ms retrieval eliminates dead air

In the STT -> retrieval -> LLM -> TTS pipeline

04

Works with Deepgram, Cartesia, OpenAI, Anthropic

And other Pipecat plugins

05

Pre-load indexes at pipeline startup

`load_index()` for fastest first-query latency

Quick Start

Explore docs
Python
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
import os
from pipecat_moss import MossRetrievalService

moss_service = MossRetrievalService(
    project_id=os.getenv("MOSS_PROJECT_ID"),
    project_key=os.getenv("MOSS_PROJECT_KEY"),
    system_prompt="Relevant passages from the Moss knowledge base:\n\n",
)

await moss_service.load_index(os.getenv("MOSS_INDEX_NAME"))

pipeline = Pipeline([
    transport.input(),
    stt,
    context_aggregator.user(),
    moss_service.query(os.getenv("MOSS_INDEX_NAME"), top_k=5),
    llm,
    tts,
    transport.output(),
    context_aggregator.assistant(),
])

Get Started in 3 Steps

01

Install pipecat-moss

Run pip install pipecat-moss to install the Pipecat pipeline processor for Moss.

02

Create and load your index

Use the Moss SDK to create_index() with your knowledge base documents, then call moss_service.load_index() at pipeline startup.

03

Insert into your pipeline

Add moss_service.query() between your STT and LLM processors. Moss retrieves context on every user turn and injects it into the LLM prompt automatically.

Frequently asked questions

Ready to ship faster AI products?

Moss gives you production-ready semantic retrieval without infrastructure complexity.

Test performanceTalk to an Engineer
Moss
AICPA SOC 2 Type 2HIPAA

Product

Founding AgentLocal Search

Use Cases

Voice AIAI CopilotsIn-App SearchOn-Device AI

Company

PricingBlogCareersBrand Kit

Resources

DocsGlossaryBenchmarks

Integrations

DSPyElevenLabsLangChainLiveKitMCP ServerNext.jsPipecatVAPIVercel AI SDKVitePress

© 2026 MOSS

Privacy PolicyTerms of ServiceTrust Center