Aside: AI memory on 80K+ devicesHow Aside powers AI memory across 80,000+ devices · 113M documents · 24.8B tokens · 150+ countries

Read moreRead the case study
Moss
usemossStart Free
Backed byY Combinator

The Local Search Engine for AI.

Semantic, keyword, and hybrid search that runs entirely on device. Embed, index, and retrieve in under 10ms with no cloud infrastructure, no vector database, and no network round trip.

Read the DocsTalk to an Engineer
Your AI app

Your AI app

Retrieval Runtime powered by Moss

Ask a question…
Moss engineidle

    <10ms

    Retrieval

    32 MB

    One Model

    100%

    Offline

    CPU Only

    No NPU Required

    Used in production by companies backed by Y-Combinator & a16z

    coderabbit logo
    aside logo
    zomato logo
    browseruse logo
    podium logo
    sigmamind logo
    zavo logo
    speko logo
    grammarly logo
    microsoft logo
    berkeley logo
    cmu logo
    stanford logo
    kindredvoice logo
    samsara logo
    rumi logo
    topmate logo
    consensus logo

    Problem

    On-device AI is here. Retrieval still needs the cloud.

    Modern AI runs on phones, laptops, browsers, vehicles, wearables, and edge devices. Retrieval still runs in the cloud, adding latency, infrastructure cost, and privacy tradeoffs.

    Cloud
    Retrieval
    Cloud round trip
    On-Device AI

    Cloud Search Infrastructure

    • Every query needs a network request
    • Vector databases become part of your architecture
    • Costs scale with usage
    • No offline support

    Traditional Embedding Pipelines

    • Large models increase app size
    • GPU and NPU compete with your AI workloads
    • Built for servers, not edge devices
    • Too heavy for phones and consumer hardware

    Local AI needs a different retrieval architecture.

    Retrieval that runs where your AI runs.

    Moss brings semantic, keyword, and hybrid search directly to where your AI runs, eliminating cloud latency, reducing infrastructure costs, and keeping data on device.

    Moss on Device AI
    • faster
    • lower cost
    • privacy

    On Device AI

    Ground local AI with instant access to user context.

    Intelligent In-App Search

    Search everything inside your application by meaning, not keywords.

    Cross-App Retrieval

    Retrieve context across files, notes, photos, messages, and application data.

    Built for Production

    Runs entirely on CPU with no cloud infrastructure or external services.

    • Messages1ms

      100K+ msgs

      SMS, chat, attachment context

    • Files1ms

      10K+ files

      Document text + filenames + metadata

    team offsite san francisco
    1. Gallery

    2. Notes

      offsite agenda — day 1, Fort Mason

    3. Messages

      …flights into SFO are booked for the offsite!

    4. Files

      Offsite plan June 20.docx

      Offsite plan June 20 (2).pdf

    • Gallery3ms

      50K+ photos

      Image captions + EXIF + people tags

    • Notes0.8ms

      100K+ notes

      Notes body + handwriting OCR

    The complete retrieval pipeline. Running locally.

    Moss runs the entire retrieval pipeline on device.
    Whether you're powering AI assistants, semantic search, RAG, recommendations, memory, or system search, it all happens locally.

    No cloud APIs.
    No embedding service.
    No vector database.
    No round trip.

    Just instant retrieval where your users already are.

    Moss engine

    1. Embedding query locally
    2. Searching memory index
    3. Running hybrid search (vector + BM25)
    4. Scanning 142,384 memories
    5. Applying RRF ranking
    6. Applying recency signals
    7. Filtering by workspace permissions
    8. 3 relevant memories retrieved
    9. last_month_investor_update
    10. mrr_chart · recently opened
    11. leads · updated yesterday
    12. Total latency: 4.8ms

    Your AI app

    finish the investor update

    I found last month's update, the MRR chart, and the Leads from yesterday.

    I'll draft this month's version with the same structure and update the metrics.

    1. Content
    2. EmbeddedOn Device
    3. IndexedOn Device
    4. RetrievedOn Device
    5. Grounded AI & Search

    One SDK. Complete local retrieval.

    Search

    Everything developers need in one SDK.

    • Semantic Search
    • Keyword Search
    • Hybrid Search
    • Cross Index Search

    Embeddings

    One lightweight model for every modality.

    • Text
    • Images
    • Audio
    • 32 MB INT8
    • CPU Execution

    Retrieval Engine

    Purpose built for edge hardware.

    • Proprietary indexing
    • Memory mapped retrieval
    • Instant indexing
    • Less than 1% recall loss after quantization

    Built for modern AI products.

    On Device AI

    Ground local AI in documents, messages, notes, photos, and app data. Nothing leaves the device.

    Consumer Apps

    Search every piece of user content inside your app. Built for productivity, media, ecommerce, travel, finance, and health.

    OEMs

    One local retrieval engine across phones, vehicles, TVs, wearables, and embedded devices. System-wide search, fast and offline.

    Enterprise Applications

    Secure AI search that keeps sensitive company knowledge on employee devices.

    Faster than cloud search.

    SystemAverage Search Latency
    Moss3.1 ms
    ChromaDB351.8 ms
    Pinecone432.6 ms
    Qdrant597.6 ms

    Benchmarks include embedding inference and retrieval across 100,000 documents. Most competing benchmarks measure retrieval only.

    The future of AI is local.
    Retrieval should be too.

    Embed. Index. Retrieve. All on the device.

    Read the DocsTalk to an Engineer
    Moss
    AICPA SOC 2 Type 2HIPAA

    Product

    Founding AgentLocal Search

    Use Cases

    Voice AIAI CopilotsIn-App SearchOn-Device AI

    Company

    PricingBlogCareersBrand Kit

    Resources

    DocsGlossaryBenchmarks

    Integrations

    DSPyElevenLabsLangChainLiveKitMCP ServerNext.jsPipecatVAPIVercel AI SDKVitePress

    © 2026 MOSS

    Privacy PolicyTerms of ServiceTrust Center