Companies turning enterprise tribal knowledge into agent-usable context
August 26 at 09:18 · $0.103 total
Sourcing Memo: Tribal Knowledge → Agent-Usable Context
1. Glean — Enterprise search + knowledge graph across SaaS apps, now exposing that graph to agents via APIs. Fits as the category anchor, though late-stage (~$7B val); useful as the "incumbent risk" benchmark rather than a target.
2. Unstructured — Converts messy enterprise docs (PDFs, emails, decks) into LLM/agent-ready structured data. Series B, strong OSS adoption; the ingestion layer of this thesis.
3. Dust — Platform for building enterprise agents grounded in company data/connections (Notion, Slack, Drive). Paris-based, Sequoia-backed, early but with real mid-market logos; founded by ex-OpenAI/Stripe.
4. Credal — Secure layer connecting internal knowledge sources to LLMs/agents with permissioning and audit. Early-stage (Spark-backed), fits the "agents need permission-aware context" angle precisely.
5. Onyx (formerly Danswer) — Open-source enterprise search/RAG over internal tools, evolving toward agent-facing knowledge APIs. Seed-stage, YC, strong OSS traction — classic bottom-up wedge.
6. Sana (Sana AI) — Knowledge platform turning company docs, meetings, and training content into an assistant/agent layer. Stockholm-based; started in learning, pivoting hard into agentic knowledge work.
7. Graphlit — API-first platform for building knowledge graphs from unstructured enterprise content for RAG/agent apps. Very early, developer-focused; a picks-and-shovels bet on context infrastructure.
8. WRITER — Enterprise genAI with a proprietary "Knowledge Graph" that grounds agents in company data/terminology. Growth stage but a signal that grounding in tribal knowledge is the enterprise buying criterion.
9. Zep — Memory and knowledge-graph layer (Graphiti OSS) for AI agents, persisting learned org/user context over time. Seed-stage; addresses the "agents forget everything" gap directly.
10. Guru — Captures verified team knowledge (cards, expert verification) and now pipes it into AI answers/agents. Older company, but its human-verification workflow is exactly the trust layer agents lack — potential acquisition comp.
(Deliberately excluded some I couldn't verify still operate independently.)
---
White-space gaps
1. Passive tribal-knowledge capture. Nearly everyone indexes what's already written down. The real tribal knowledge lives in Slack threads, Zoom calls, code review comments, and hallway decisions — nobody has nailed ambient capture → structured, permissioned context with human verification in the loop. Guru's verification model + Zep's memory model, but passive.
2. Context provenance & decay. No one is building the "freshness and trust layer": which context is stale, who vouched for it, what an agent is allowed to assert vs. must caveat. As agents act (not just answer), liability for acting on outdated tribal knowledge becomes a budgeted enterprise problem — and there's no clear owner of that layer today.
Here are 10 early-stage companies fitting the thesis of turning unstructured enterprise tribal knowledge into agent-usable context, focusing on lesser-known and highly technical startups:
1. Reducto — An API that converts complex, unstructured enterprise documents (PDFs, nested spreadsheets, slide decks) into clean, chunked data specifically formatted for LLM and agent ingestion. Why it fits: YC W24; strong early signal from AI engineers who are blocked by the inability of standard OCR to accurately extract tables, graphs, and spatial context from legacy enterprise files.
2. Greptile — An API that ingests codebases, pull requests, Jira tickets, and Slack discussions to give AI agents deep, contextual understanding of software projects. Why it fits: YC W24; high traction among devtool startups building agentic features that need to understand undocumented engineering decisions and tribal codebase history.
3. Mem0 (formerly Embedchain) — A memory layer for AI applications that automatically extracts, structures, and manages user preferences and enterprise context across continuous sessions. Why it fits: YC S24; massive open-source traction (20k+ GitHub stars) signaling strong developer demand for plug-and-play, cross-session agent memory.
4. Zep — A long-term memory and context API for AI agents that extracts facts, summaries, and relational knowledge graphs from historical conversational data. Why it fits: Seed stage; backed by prominent AI angels, it solves the "agent amnesia" problem by turning ephemeral chat logs and customer interactions into structured, queryable context.
5. Cassidy — A platform that connects directly to enterprise tools (Slack, Notion, Drive) to build a unified, continuously updated context layer for deploying custom, role-specific AI agents. Why it fits: Seed stage; gaining rapid organic adoption in mid-market companies looking for a "Glean for agents" without the massive enterprise deployment overhead.
6. Nuclia — An end-to-end API that automatically transcribes and indexes unstructured data from any internal source—crucially including video and audio—into an optimized vector database for agents. Why it fits: Seed/Early Series A; strong European footprint and unique technical moat in turning multimedia tribal knowledge (e.g., recorded Zoom meetings, Loom tutorials) into text-searchable agent context.
7. Dashworks — Originally an enterprise search tool, now an infrastructure layer that creates an automated, permission-aware knowledge graph from company SaaS apps for AI agents to query. Why it fits: Series A (Sequoia/YC); successfully pivoting their deep integration stack into an agentic infrastructure narrative, leveraging their existing ability to map complex enterprise permissions.
8. Khoj — An open-source AI workspace that connects to local and cloud-based notes, docs, and chats to serve as a contextual backend for AI assistants. Why it fits: YC W24; strong open-source community signal from privacy-conscious enterprises (defense, healthcare, legal) that require self-hosted, air-gapped context engines.
9. Klu — A data platform that unifies internal enterprise data sources so teams can design, evaluate, and deploy LLM applications strictly grounded in company-specific context. Why it fits: Pre-seed/Seed; highly focused on the "data-to-agent" pipeline, showing strong early signal by bridging the gap between raw internal data and LLM prompt engineering/evaluations.
10. Quivr — An open-source RAG framework that acts as a generative "second brain" for enterprises, turning fragmented files and chat histories into an API-accessible knowledge base. Why it fits: Seed stage; explosive GitHub growth (30k+ stars) proving massive organic developer appetite for open-source, customizable enterprise context layers over closed-source alternatives.
***
The 2 Most Interesting White-Space Gaps
1. Implicit Workflow Extraction (Action-to-Context) Current startups focus heavily on text-based tribal knowledge (docs, code, Slack). However, massive amounts of enterprise knowledge are implicit—living in how an employee navigates a CRM, clicks through an ERP, or triages a dashboard. There is a massive gap for companies using multimodal models (like computer vision/screen recording) to observe human workflows and automatically compile them into structured Standard Operating Procedures (SOPs) that an autonomous agent can execute.
2. Dynamic Identity and Permission Mirroring (RBAC for Agents) Turning tribal knowledge into vector embeddings is becoming commoditized, but doing so while strictly respecting enterprise permissions is not. If an agent searches a company's knowledge base on behalf of a junior employee, it must not retrieve context from a Slack thread involving the CEO and HR. Building a universal identity layer that perfectly mirrors complex, constantly changing Role-Based Access Control (RBAC) across Jira, Google Drive, and Slack specifically for autonomous agents remains a largely unsolved infrastructure gap for early-stage builders.