Arango logo

What’s new in the data platform

Features and improvements released for the Contextual Data Platform

v4.1.1-preview (September 2026)

Limited Release

This preview release is available to selected customers as part of a Limited Release. It is a maintenance release with bug fixes and security improvements.

MLflow

Agentic AI Suite

Authentication is now enforced on the route of the integrated MLflow service. As a result, the MLflow web interface is only available through the Arango Contextual Data Platform web interface, under AI Tools. Opening https://<EXTERNAL_ENDPOINT>:8529/mlflow/ directly in a browser is no longer possible.

Programmatic access is unchanged. The official MLflow client and other HTTP callers continue to use the same endpoint with a valid JWT, see Programmatic access.

AutoGraph

Agentic AI Suite

  • More accurate answers: Retrieval no longer matches questions against the wrong passages, content from your documents is no longer left out of the graph along with its connections, and search stays within the boundaries of each project.
  • More reliable builds: Large document sets and large file uploads no longer cause timeouts or failures, builds no longer fail while a valid API key is in use, and existing Corpus Graphs can be updated after they have been built.
  • Visible failure causes: A failed Knowledge Graph build now shows the strategy execution summary, so you can see why it failed.
The fixes for wrong passages and missing content only apply to new imports. Graphs built before this release keep the old behavior. If your graph was built from more than one document, import the content again to get the corrected result.

GraphML

Agentic AI Suite

Prediction jobs no longer remain in the Pending state after featurization, training, and model generation have finished successfully.

v4.1.0 (August 2026)

AutoGraph Studio

Agentic AI Suite

The AutoGraph and GraphRAG web interfaces have been unified into AutoGraph Studio, a single workflow that covers document upload, model configuration, corpus and knowledge graph building, retriever deployment, and querying your Context Graph.

The workflow has two stages, and the first one can be your finish line:

  • AutoGraph: Analyzes your documents, builds the Corpus Graph, and generates the strategies for the Knowledge Graph. If all you need is the Context Graph, you can stop here and explore it in the Graph Visualizer or query it directly.
  • AutoRAG: Optionally deploys retrievers on top of that Context Graph, so your agents and applications can ask questions against it.

The terminology has been aligned across the documentation: Corpus Graph, Knowledge Graph, and Context Graph now refer to distinct artifacts of the AutoGraph pipeline. The standalone GraphRAG web interface has been removed; use AutoGraph and the new AutoGraph Studio web interface instead. You can also use the Importer and AutoRAG APIs.

Incremental Graph Updates

Agentic AI Suite

Incremental Graph Updates keep a Knowledge Graph current after it has been built. It is faster and more efficient than to recreate it, and can significantly reduce the LLM cost and latency.

You can insert new documents, delete obsolete ones, and replace changed ones without re-running the corpus build, the RAG Strategizer, and a full orchestration pass. Existing clusters and strategy profiles are preserved, and only the documents that actually changed are processed.

AutoGraph also tracks how far a FullGraphRAG partition has drifted since it was last clustered and flags it, so you can trigger reclustering when you choose to. Reclustering is never automatic.

In this release, incremental updates are available through the HTTP API only.

File Parser Service

Agentic AI Suite

A new internal service for converting documents has been added. Both AutoGraph and the Importer now delegate the conversion of documents into Markdown to the new File Parser service.

PDF files, including scanned documents that are read using OCR, and some Office document formats (.docx, .pptx, .doc, .ppt) are now officially supported. The service additionally extracts embedded images together with their surrounding text (if requested), so that the Importer can pick them up as semantic units. For what each format guarantees, see Document conversion and supported formats.

The new service is designed for horizontal scalability, using two worker tiers, one for PDF documents and one for everything else, with 3 worker pods per tier by default. Deployments in AMP run these defaults unchanged. For self-hosted clusters, see Tuning the File Parser.

File Manager

Platform Suite

  • RAG input files are now organized into scopes, an ordered list of up to five labels that addresses a file within a database. Each service maps its own concepts, such as projects and modules, onto scope levels. A file is identified by database, scope, and name, and re-uploading the same name into the same scope creates a new version.
  • You can attach custom metadata to uploaded files. The reserved citable_url key lets AutoRAG resolve citations back to the original source document.
  • You can upload many files in one request, up to 100 files and 2 GiB, either with a manifest that places each file individually or with one shared scope. The response reports the outcome per file, so a batch in which single files fail still stores the rest.
  • Files can be browsed as a folder tree: a scope reports its child scopes with their file counts along with the files that sit directly in it. Listing files accepts a scope, which covers everything below it, and a case-insensitive name search, and it returns the latest version of each file instead of the full history.
  • Files can be locked against deletion, individually, in bulk, or for a whole scope. Every delete path skips a locked file and reports it instead of removing it, which is how AutoGraph protects the files that a corpus still references.
  • Files can be deleted in bulk by id or by the scope that holds them.

For the endpoints and the status code changes, see API changes.

Container Manager

Platform Suite

  • Services that serve their own HTML interface can be registered as Apps. An App appears in the platform’s Apps catalog and is rendered embedded in the web interface, which is useful for custom dashboards, admin panels, and interactive tools next to your data.
  • Node.js 22 is now available as a base image (node22base) for code-based deployments via the API, alongside the Python variants.

v4.0.2 (May 2026)

This is a maintenance release.

Container Manager

Platform Suite

The Container Manager base images (base, PyTorch, and cuGraph variants) have been updated to Python 3.12; service packages must now target Python 3.12. The release also includes security fixes.

v4.0.1 (May 2026)

This release contains improvements and refinements to features introduced in v4.0.0.

AutoGraph

Agentic AI Suite

  • Corpus build failures caused by the LLM or embedding provider now surface a machine-readable error_code on the build status response (authentication failed, permission denied, rate limited, quota exceeded, or API key missing), so clients can react to each case instead of parsing free-text messages. The error reference also adds HTTP 429 (provider rate-limited or quota exhausted), expands the meanings of 401 and 403 to cover LLM provider auth and permission failures, and explains why an accepted (202) async job can still fail later.
  • The new Known Limitations section in the error reference documents two important behaviors: citation extraction and SemanticUnits linking are not yet automatic (you provide citable_url and run your own post-processing); and VectorRAG partitions cannot serve Global or Local queries because they skip entity and community extraction.
  • The RAG Strategizer response is documented more precisely: rag_partition_id suffixes (_a = FullGraphRAG, _b = VectorRAG), the full list of FullGraphRAG importer tunables returned in parameters, empty entity_types for VectorRAG clusters, and an entity_generation_error field that appears when LLM-driven entity-type generation fails for a cluster.

Importer

Agentic AI Suite

  • The Importer now auto-detects each chat model’s context window and picks a sensible completion-token cap, so common OpenAI models (GPT-4o, GPT-4 Turbo, GPT-5.4 Nano, o1, o3) work without manual tuning. New environment variables (CHAT_MAX_COMPLETION_TOKENS, CHAT_MODEL_CONTEXT_TOKENS, GRAPHRAG_LLM_PROMPT_TOKEN_BUDGET) let you override the defaults for private fine-tunes or sparse graphs. See LLM configuration.
  • Before each chat call, the Importer re-tokenizes the actual prompt and truncates it if it would exceed the model’s window, so jobs no longer fail with context_length_exceeded on long prompts.
  • The Importer now auto-detects newer OpenAI models that require /v1/responses (for example gpt-5.4-pro, o3-pro), retries the call via the Responses API, and caches the result so subsequent calls skip the failing chat-completion attempt.
  • Provider errors during graph build are mapped to short remediation messages (insufficient quota, invalid API key, rate limit, timeout, 5xx, context length exceeded) and stored on the service status, so operators see actionable text instead of raw SDK output.
  • The model used for image description during semantic-unit processing is now configurable via the MULTIMODAL_MODEL environment variable (default gpt-4o-mini), and it honors the same token budget and Responses API settings as the rest of the pipeline.

Retriever (now AutoRAG)

Agentic AI Suite

  • Response caching (use_cache: true) now works for every query type (GLOBAL, LOCAL, UNIFIED, and CUSTOM); previously only some query types could be cached.
  • show_citations is documented as a no-op in Deep Search (use_llm_planner=true) and GLOBAL queries, because those modes always strip citations regardless of the flag. The parameter still applies to LOCAL, UNIFIED, and CUSTOM queries.
  • For CUSTOM queries, an individual tool’s own show_citations: false configuration can suppress citations from that tool’s results even when the request-level flag is true, so you can mix citation behavior across the components of a custom retriever.

Default AI models

Agentic AI Suite

Default OpenAI chat model upgraded from gpt-4o to the GPT-5.4 family. The Importer and Retriever now default to gpt-5.4-nano; the Natural Language to AQL service (AQLizer) defaults to gpt-5.4. Ada also adds Anthropic, OpenRouter, and Custom Endpoint as provider options alongside OpenAI.

MLflow

Agentic AI Suite

The integrated MLflow service has been upgraded to MLflow 3.x.

License activation

A new web-based License Activation portal is available for internet-connected deployments as an alternative to the Platform CLI, with Managed, Inventory, and Generic modes.

v4.0.0 (April 2026, General Availability)

The Arango Contextual Data Platform has been officially released and comes with various improvements and major additions to the unified web interface, the Platform Suite and the Agentic AI Suite.

The minimum required ArangoDB version is the Enterprise Edition v3.12.9.

AutoGraph

Agentic AI Suite

AutoGraph is an automation copilot that analyzes enterprise documents, discovers natural knowledge domains, and automatically builds optimized knowledge graphs for intelligent retrieval at scale.

Key features:

  • Automated domain discovery: Analyzes document relationships and discovers natural clusters using graph algorithms, creating specialized RAG partitions per domain.
  • Intelligent RAG strategy selection: Automatically recommends FullGraphRAG or VectorRAG for each domain based on content complexity.
  • Knowledge graph creation: Transforms documents into structured knowledge graphs with entities, relationships, and semantic connections.
  • Natural language querying: Chat with your knowledge graph using natural language to ask questions and retrieve insights from your documents.
  • Web interface: Streamlined workflow guides you through document upload, corpus building, strategy generation, knowledge graph import, retriever deployment, and chat.

Ada

Agentic AI Suite Beta

Ada is a new AI digital assistant integrated into the Arango Contextual Data Platform. It lets you interact with your database using natural language, generate and execute AQL queries, explore collections and data structures, and save reusable query artifacts through a conversational chat interface.

Graph Analytics

Agentic AI Suite

Graph Analytics now includes a web interface and a new AQL-based data loading API.

  • Web interface: A new graphical interface is now available for managing Graph Analytics Engines, loading graphs, running algorithms (PageRank, Connected Components, Label Propagation, and more), and monitoring job progress. The interface provides an intuitive workflow for engine management, graph loading, algorithm execution, job monitoring, and trigger jobs to persist results of algorithms into collections.

  • AQL-based data loading API: A new API endpoint allows you to import graph data using custom AQL queries. Queries are organized into phases that run sequentially, while queries within each phase execute in parallel for optimal performance.

Graph Visualizer

Platform Suite

The Graph Visualizer now supports exporting graph data in CSV format. Export the entire visible canvas, or export nodes and edges separately with all document attributes included.

Query Editor

Platform Suite

The Query Editor has been extended with the following capabilities:

  • Graph visualization: If a query returns edges or traversal paths, the results are shown by an embedded graph visualizer. You can still switch to a JSON view mode of the results.

  • Download results: You can download the results of queries in JSON and CSV format.

  • Syntax highlighting: AQL queries in the query editor are colorized for better readability.

AQL Optimizer (Reasoner)

Agentic AI Suite Beta

A new Optimize button has been added to query tabs for AI-powered query optimization. The Reasoner analyzes your AQL query and suggests improvements through a streaming chat interface with real-time tool call and validation feedback. This feature requires a license.

Container Manager

Platform Suite

The Container Manager enables you to deploy and manage custom services directly within the Arango Contextual Data Platform, running your own applications alongside platform services.

Deploy services by uploading source code packages (.tar.gz) or providing Docker image URLs, with support for Python 3.12 runtimes (including PyTorch and cuGraph variants). Services can be scoped globally or per-database, with version management and deployment via web interface or API.

File Manager

Platform Suite

The File Manager provides a centralized interface for viewing and managing data stored by platform services, including container service files, RAG content, and AutoGraph files.

Secrets Manager

Platform Suite

The Secrets Manager has been introduced for managing API keys and credentials across the platform. Secrets are encrypted at rest and accessible to services via sidecar containers, with support for bulk operations, import/export, and multiple secret types.

Monitoring

Platform Suite

Integrated Monitoring with Grafana and Prometheus provides observability for the entire deployment. Both tools are embedded in the unified web interface with authenticated access for tracking performance metrics, cluster health, and resource utilization.

Dashboard

Platform Suite

A new home screen has been added, providing the following information and actions:

  • A cluster respectively single server health overview
  • Shard distribution information and rebalancing options
  • The cluster maintenance status

Cypher to AQL (experimental)

Platform Suite

An experimental service for translating Cypher queries to ArangoDB’s query language AQL has been added. The arango-cypher2aql service provides an API for parser-based translation so that you can reuse existing Cypher knowledge.

v3.0.0 pre-release 2 (October 2025)

This release includes new features and enhancements for the Contextual Data Platform web interface as well as the components of the Agentic AI Suite.

The minimum required ArangoDB version has been raised to Enterprise Edition v3.12.6.

GraphRAG enhancements

Agentic AI Suite

  • Instant and Deep Search: New Retriever (now AutoRAG) search methods optimized for different use cases. Instant Search provides fast responses with streaming support. Deep Search offers detailed, accurate responses for complex queries requiring high accuracy. Both methods are accessible via the API or the AutoGraph web interface.

  • Update Knowledge Graphs: Add additional data sources to existing Knowledge Graphs through the web interface. Upload new files to automatically update the Knowledge Graph and underlying collections with new data.

  • Unified LLM provider configuration: Simplified deployment configuration using OpenAI-compatible APIs. Mix and match providers for chat and embeddings (e.g., use OpenRouter for chat and OpenAI for embeddings). Support for OpenAI, OpenRouter, Gemini, Anthropic, and any self-hosted LLM with OpenAI-compatible endpoints.

AQLizer

Agentic AI Suite Beta

The Natural Language to AQL Translation Service enables you to query your ArangoDB database using natural language or get LLM-powered answers to general questions.

You can generate AQL queries from natural language directly in the Query Editor using the AQLizer mode. More advanced features are available via the API.

Query Editor

Platform Suite

A new Query Editor has been integrated into the Arango Contextual Data Platform web interface for writing, executing, and managing AQL queries.

Key features:

  • Tabbed interface: Work on multiple queries concurrently with side-by-side query and results views.
  • Query operations: Run, explain, and profile queries with dedicated buttons and result history tracking.
  • Saved queries: Save and share frequently used queries with all users in the database, persisted across sessions.
  • Query monitoring: View running queries and slow query logs, with the ability to kill long-running operations.
  • Flexible viewport: Drag and drop tabs to reorganize panels horizontally or vertically.

Graph Visualizer enhancements

Platform Suite

The Graph Visualizer has been significantly enhanced with new visual customization capabilities, improved navigation features, and better performance for exploring large-scale graphs.

Key improvements:

  • Icon assignment: Assign pictograms to node collections for quick visual identification of entity types on the canvas.
  • Theme support: Create and manage multiple themes to highlight different aspects of graph data, with default themes that automatically color different collections on the canvas.
  • Shortest path: Find and visualize the shortest path between two selected nodes directly on the canvas.
  • Enhanced tooltips: Hover over nodes and edges to view document IDs and customizable additional attributes without opening the full properties dialog.
  • Bulk selection: Select all nodes or edges of a specific type (collection) from the Legend panel, showing the count of elements per collection.
  • Edge properties view: View and edit edge properties through a dedicated properties dialog with Form and JSON editing modes.
  • Attribute-based styling: Define conditional styling rules based on document attributes to dynamically color and style nodes and edges (e.g., apply colors based on genre or other field values).
  • Performance improvements: Optimized rendering for large graphs with millions of nodes and edges.

Platform Services

Platform Suite

  • Secrets Manager: Store secrets like API keys for Large Language Model (LLM) for easy use across the Contextual Data Platform. Secrets are encrypted at rest and can be accessed by services via a metadata sidecar container.

v3.0.0 pre-release 1 (July 2025)

This release marks the initial internal launch of the Arango Contextual Data Platform and its Agentic AI Suite and Platform Suite.

The minimum required ArangoDB version is the Enterprise Edition v3.12.5.

Introducing the Arango Agentic AI Suite

Agentic AI Suite

What’s included:

  • GraphRAG: Transform unstructured documents into intelligent knowledge graphs and natural language querying through the Importer and Retriever (now AutoRAG) services.

  • GraphML: Apply machine learning to graphs with node classification and embedding generation, built on GraphSAGE framework.

  • Graph Analytics: Run algorithms like PageRank, Connected Components, and more.

  • Jupyter Notebooks: Launch integrated Jupyter notebook servers with pre-installed ArangoDB drivers and data science libraries for interactive experimentation.

  • MLflow Integration: Use MLflow as a model registry for private LLMs and machine learning experiment tracking.

  • Triton Inference Server: Host private Large Language Models using NVIDIA Triton Inference Server for secure, on-premises AI capabilities.

Introducing the Arango Platform Suite

Platform Suite

What’s included:

  • ArangoDB Enterprise Edition: Multi-model database foundation supporting graphs, documents, key-value, vector search, and full-text search capabilities.
  • Unified web interface: Single interface for accessing all Contextual Data Platform services and components.
  • Graph Visualizer: Sophisticated web-based interface for interactive graph exploration, visual customization, and direct graph editing.
  • Query Editor: Write, run, and analyze AQL queries using an IDE-like interface with tabs, result history, query management, and more.
  • Kubernetes orchestration: Powered by the official ArangoDB Kubernetes Operator for automated deployment, scaling, and management.
  • Operational features: Enterprise-grade features including high availability and monitoring, comprehensive APIs and connectors, and centralized orchestration and resource management.
  • Additional services: Cypher2AQL service for translating graph queries written in Neo4j’s Cypher query language to ArangoDB’s AQL query language (experimental)