Arango logo

AutoGraph

AutoGraph structures enterprise data into a Context Graph with domain-aware retrieval strategies, and AutoRAG deploys the retrievers that answer questions from it, providing AI copilots and agents with production-grade context infrastructure

STAGE 1 · AUTOGRAPH — build the Context GraphSTAGE 2 · AUTORAG — answer questions from itYour documentsPDF, Office, Text,MarkdownAutoGraphclusters documents,assigns strategiesContext GraphCorpus Graphtopic domainsKnowledge Graphentities + embeddingsAutoRAG retrieverfinds the relevantcontextLLM providergenerates answergrounded answerUser / Agentasks a questionquestion
AutoGraph builds the Context Graph from your documents; AutoRAG serves questions from it. View file

The workflow has two stages, both driven from the AutoGraph Studio view of the web interface:

  1. AutoGraph builds your Context Graph — the Corpus Graph that maps your documents into topic domains, plus the Knowledge Graph built from them. This can be your finish line.
  2. AutoRAG is the retrieval layer on top. It deploys the retrievers that answer questions from that Context Graph, so your agents and applications can query it.

What is AutoGraph?

AutoGraph is a large-scale RAG system that delivers strong accuracy at the quality-cost tradeoff you choose. It supports benchmarking, testset creation, automated ontologies, and extensibility to new RAG methods - distilling lessons learned from running RAG at some of the world’s largest enterprises.

Under the hood, AutoGraph is an automation copilot that analyzes enterprise documents, discovers natural knowledge domains, and builds semantic infrastructure for intelligent retrieval at scale - importing documents, generating embeddings, building knowledge graphs, assigning RAG strategies per domain, and orchestrating downstream GraphRAG builds.

Think of it as a self-organizing knowledge system. Instead of manually categorizing documents or designing taxonomies, AutoGraph handles the following:

  1. Analyzes document relationships automatically
  2. Discovers natural domain clusters using graph algorithms
  3. Creates specialized RAG partitions per domain
  4. Optimizes retrieval strategies per domain
  5. Routes queries intelligently to relevant domains

The result is a domain-aware knowledge base that scales horizontally across machines.

Why AutoGraph?

AutoGraph automatically discovers that enterprise data naturally divides into knowledge domains, with each domain deserving its own optimized processing and retrieval strategy. By building a Corpus Graph (the map of your knowledge) and importing each domain into specialized RAG partitions, AutoGraph enables:

  • Automatic domain discovery
  • Horizontal scaling across machines
  • Cost-optimized processing
  • Intelligent retrieval

This approach solves the compounding challenges modern enterprises face:

  • Fragmentation: Unifies data scattered across dozens of systems into a connected knowledge graph
  • Scale: Handles thousands to millions of documents through horizontal scaling
  • Heterogeneity: Processes simple FAQs differently from complex technical specs
  • Cost: Matches processing intensity to content complexity, avoiding expensive LLM waste
  • Performance: Searches only relevant domain partitions instead of the entire corpus
  • Change: Lets you add, remove, and replace individual documents instead of rebuilding the corpus every time a few files change

Traditional RAG solutions treat all documents the same way, leading to either inadequate processing of complex content or wasteful over-processing of simple content. AutoGraph adapts to your data.

By organizing enterprise data into contextual knowledge graphs, AutoGraph creates a semantic data layer that represents relationships between business entities, systems, and operational events. This enables AI agents to:

  • Reason across enterprise relationships
  • Understand real-time operational states
  • Operate within governance policies
  • Produce explainable outputs with traceable lineage

RAG Strategizer

Not all content is equally complex. The RAG Strategizer examines the domain clusters in the Corpus Graph and assigns each one a processing strategy: complex domains get a full knowledge graph with extracted entities and relationships (FullGraphRAG, labeled GraphRAG in the web interface); simpler domains get a lighter partition that skips entity extraction (VectorRAG). For FullGraphRAG domains, it also generates a domain-specific ontology (the entity types to extract), so the resulting knowledge graph reflects the concepts that actually matter in that content.

Incremental Graph Updates

Document sets change over time. Contracts are amended, specifications are revised, and obsolete files need to be removed. If you rebuild the whole corpus for a few changed documents, you pay again for the extraction, the embeddings, the clustering, and a full Importer run.

Incremental Graph Updates keep a knowledge graph up-to-date after it has been built. You insert, delete, or replace individual documents, and only what actually changed is processed. Existing clusters and strategy profiles are kept, and a new document joins the cluster closest to it, so the whole domain does not have to be clustered again. AutoGraph also measures how far each partition has drifted since it was last clustered and flags the ones that may need a refresh, but it never reclusters on its own. That decision, and the cost of it, is up to you.

See Incremental Graph Updates.

What’s next

  • Concepts: Knowledge graphs, LLMs, and the GraphRAG approach that AutoGraph builds on.
  • Quick Start: Turn a pile of documents into a knowledge base you can chat with, with answers cited back to the source.
  • Use Cases: Understand the business value through real-world enterprise scenarios and how AutoGraph compares to traditional RAG.
  • Setup: Set up AutoGraph using the web interface or the HTTP API.
  • Web Interface: Run the complete workflow in AutoGraph Studio, from building the Context Graph to deploying AutoRAG retrievers and asking questions against it.
  • Architecture: Explore AutoGraph’s three-layer knowledge graph architecture and ArangoDB collections.
  • Design Guide: Learn how to structure your data with categories, layers, and components.
  • Incremental Graph Updates: Insert, delete, and update individual documents in a knowledge graph that has already been built.
  • API Reference: Dive into the corpus build, embeddings, RAG Strategizer, and orchestration endpoints.