Enterprise AI Agent Memory Architecture: A Selection Map from Oracle to AWS, GCP, and Open Source
AI動向 業界ニュース 41 min read

Enterprise AI Agent Memory Architecture: A Selection Map from Oracle to AWS, GCP, and Open Source

Once AI agents take on multi-day tasks, memory becomes infrastructure. This article covers Oracle AI Agent Memory's converged-database route, TiDB Vector plus Mem0's distributed HTAP route, the managed AWS Bedrock AgentCore Memory and GCP Vertex AI Memory Bank services, and four open-source frameworks: Mem0, Letta, Zep/Graphiti, and LangMem. It closes with selection guidance for regulated industries, cloud-native apps, multi-tenant SaaS, and temporal-logic use cases, plus three future trends.

Enterprise AI Agent Memory Architecture: From Oracle's Converged Database to AWS, GCP, and Open-Source Frameworks

When Agents Move from "Answering" to "Executing," Memory Becomes the Bottleneck

Large language models are undergoing a role change: from one-shot, stateless question answering to autonomous AI agents that can execute complex, long-horizon tasks. Once an agent has to keep track of the same matter for days or even weeks, "not remembering" stops being a mere experience problem and becomes an infrastructure problem that determines whether the system can exist at all.

Conventional prompt stitching and simple vector-retrieval RAG work well enough in single-turn scenarios. But in cross-session, long-running tasks they quickly expose four problems: context gets diluted by irrelevant content, retrieval latency balloons as data grows, data is repeatedly synchronized across multiple stores and becomes redundant, and there is no rigorous data governance or compliance boundary.

Around these bottlenecks, major cloud providers and the open-source community have formed two clearly different technical routes:

  • Converged Database Core: coupling agent state directly into the underlying enterprise database. Oracle AI Agent Memory is the representative.
  • Distributed middleware and dedicated storage: managing memory through an independent middleware layer with a variety of backend stores. The representatives are AWS Bedrock AgentCore Memory, GCP Vertex AI Memory Bank, and open-source frameworks such as Mem0, Letta, Zep/Graphiti, and LangMem.

Below I break down the architecture, implementation mechanisms, and applicable boundaries of both routes, then close with a selection reference for multi-cloud environments.

Oracle AI Agent Memory: Converging Memory into the Enterprise Database

From a "patchwork architecture" to a single engine

Traditional agent memory systems are usually a scattered patchwork: a vector database stores semantic fragments, a graph database maintains entity relationships, JSON stores conversation history and user settings, and a relational database handles business transactions. In enterprise settings this fragmented architecture creates three cascading problems: high operational complexity, data pipelines that are duplicated and hard to maintain, and permission and security-compliance boundaries that are hard to unify.

Oracle takes a completely different route. It positions Oracle AI Agent Memory (running on Oracle Database 23ai / Oracle AI Database) as the persistent memory core of an enterprise AI system and a second system of record. Native converged-database capabilities let a single engine support four data models at once: Vector Search, JSON Document, Property Graph, and Relational.

For developers, the entry point is a Python SDK, and the persistent memory layer is built directly on Oracle AI Database. There is no need to move memory data to an external service and sync it back. Memory data shares the same storage engine and indexing mechanisms as existing business data, so it delivers low-latency context retrieval and strong consistency guarantees at the same time.

Two kinds of memory and lifecycle governance

Oracle divides agent memory into two clear categories:

  • Working Memory: preserves task context, interaction state, short-term reasoning chains, and intermediate summaries across sessions, so an agent can resume a task precisely after an interruption.
  • Long-term Factual Memory: stores user preferences, business rules, historical task results, and facts that evolve over time.

At the retrieval layer it does more than vector similarity: it combines semantic similarity, keyword precision, metadata filtering, record type, and explicit memory scope to raise retrieval accuracy in complex business scenarios.

Governance is a key differentiator in the enterprise market. The system supports explicit scoping of memory fragments into User, Agent, and Thread levels. When business conditions change or a user exercises the "right to be forgotten," it performs cascading deletion and expiration, leaving no orphaned invalid fragments behind. These capabilities inherit enterprise security policies such as row-level security and KMS encryption. Oracle AI Agent Memory is also deeply integrated with Oracle AI Database Private Agent Factory: when defining an agent, you can configure a memory retention policy and specify the backend storage logic for session state and summaries.

TiDB Vector + Mem0: Trading for Cloud Neutrality with Distributed HTAP

Where TiDB Vector sits in the memory stack

For companies that want cloud neutrality, high-concurrency scaling, and familiarity with the MySQL ecosystem, extending vector search on the distributed HTAP database TiDB (TiDB Vector) is a pragmatic route.

The weakness of a pure vector database is that it struggles to handle complex metadata filtering efficiently and lacks strongly consistent transactional updates. Yet agent memory records are never just vectors: each record carries a large amount of relational data such as user ID, session timestamp, tenant isolation tags, and access-control columns. TiDB Vector puts high-concurrency distributed SQL transactions and high-dimensional vector search into the same system, and its distributed architecture natively supports horizontal scaling of large multi-tenant agent memory.

Mem0 extracts, TiDB persists

In practice, TiDB Vector usually serves as the backend persistent store, paired with a lightweight memory middleware such as Mem0. Mem0 provides an "Extract-and-Retrieve" style API: when a new conversation arrives, Mem0 internally calls an LLM to extract structured facts from the text, then converts those facts into vectors and JSON structures. Once TiDB Vector is configured as Mem0's backend, Mem0's CRUD (create, read, update, delete) operations map directly onto distributed vector and relational SQL operations in TiDB.

The value of this combination has two layers. The first is multi-cloud portability: whether the agent runs on AWS EKS, Google Cloud GKE, or TiDB Cloud, the memory-access codebase can stay unified. The second is real-time analytics: thanks to HTAP, an analytics system can run real-time SQL aggregations directly over massive historical memory data—say, analyzing high-frequency user-preference trends—without affecting the low-latency responses of online agent memory retrieval.

Managed Agent Memory on AWS and GCP

For companies deeply tied to a public-cloud ecosystem, AWS and Google Cloud both offer platform-level managed agent memory services aimed at lowering the developer's infrastructure burden and delivering out-of-the-box context persistence.

Amazon Bedrock AgentCore Memory

AWS folds agent memory into the Amazon Bedrock AgentCore ecosystem (which includes AgentCore Runtime, AgentCore Memory, and AgentCore Observability). Its design cleanly separates two concepts: the isolated session at runtime, and persistent long-term memory.

Runtime sessions are maintained by AgentCore Runtime and come with a strict session-timeout mechanism: 15 minutes by default, extendable up to 8 hours. Raw event data is automatically cleaned up after the session ends or once it exceeds the TTL (typically 30 to 90 days). Long-term memory is provided by AgentCore Memory, which uses asynchronous extraction and consolidation: after session events occur, a managed background process asynchronously calls a large model and, following predefined Memory Strategies, extracts structured insights from the raw conversation stream.

AWS ships three standard long-term memory strategies:

  • Semantic Strategy: extracts and stores the facts and knowledge mentioned in the conversation, for example recording a customer company's size and cross-region office layout.
  • Summary Strategy: maintains a running conversation summary across sessions, capturing key decisions and process progress.
  • User Preferences Strategy: automatically extracts a user's behavioral preferences, language style, or coding conventions.

On cost, the main expense of AgentCore Memory comes from Bedrock LLM calls during memory consolidation. For systems with a large existing memory footprint, AWS recommends relevance-threshold-based pruning and scheduled merges to avoid a surge in background token consumption. On multi-tenant security, developers must force a binding between the Session and the Authenticated Owner at the application layer to prevent cross-tenant, unauthorized memory retrieval.

GCP Vertex AI Agent Engine Memory Bank

Google Cloud offers the Memory Bank managed service within Vertex AI Agent Engine, working mainly with Google ADK (Agent Development Kit) and Vertex AI Sessions.

Its core idea is dynamic fact extraction plus intelligent deduplication and consolidation. The mechanism: whenever an agent and a user finish multiple rounds of conversation in Vertex AI Sessions, the session data is submitted to Memory Bank. The analysis engine automatically extracts structured facts (such as identity, dietary restrictions, or room preferences) and compares them against existing memory under that user ID. Memory Bank does not mechanically append duplicates; it automatically performs consolidation and updates.

Retrieval supports two paradigms. The first is Scope-Based Retrieval, which pulls all memory records for a specific user to build a global profile—suitable for admin backends or full-context initialization. The second is Similarity Search, which uses vector semantic matching to quickly extract relevant context for a specific question—suitable for high-concurrency, real-time conversation.

Comparing the four options

Dimension / PlatformOracle AI Agent MemoryAWS Bedrock AgentCore MemoryGCP Vertex AI Memory BankTiDB Vector + Mem0
Underlying storage dependencyOracle AI Database 23ai (converged engine)Managed AWS storage + Bedrock modelsBuilt-in storage in Vertex AI Agent EngineTiDB distributed database (MySQL protocol)
Data consolidation mechanismNative multi-model correlation and retrieval inside the databaseAsynchronous background LLM extraction (semantic/summary/preferences)Session-based extraction with intelligent dedup and consolidationDriven by the Mem0 extraction layer, written into distributed SQL
Multi-cloud and deployment flexibilityBest on OCI; supports OCI Dedicated/CustomerLimited to AWS cloud-native environmentsLimited to the GCP Vertex AI ecosystemCloud-agnostic (AWS, GCP, self-managed K8s)
Lifecycle controlCascading deletion, explicit User/Agent/Thread scopesAutomatic TTL expiration, pruning strategies, and GuardrailsAutomatic consolidation and dedup per User IDExplicit CRUD API control, reliant on database TTL
Main cost driversDB instance/service resource capacityBedrock LLM asynchronous consolidation tokensAgent Engine managed API and inference feesTiDB node compute/storage resources

Four Open-Source Memory Architectures for Cloud Neutrality

If you want to build agent systems flexibly on AWS, GCP, or a private cloud, the four most representative open-source memory architectures are Mem0, Letta (formerly MemGPT), Zep/Graphiti, and LangMem. They differ fundamentally in data structure, update mechanism, and temporal-reasoning capability.

Mem0: lightweight fact extraction and vector memory

Mem0's design philosophy is to offer a minimal "drop-in" memory API. It avoids complex operational-framework patterns and focuses on extracting atomic facts from the conversation stream and binding them to a specific scope (User, Agent, Session, or App).

When a new conversation arrives, Mem0 calls an LLM to distill short facts and converts them into embeddings stored in a vector database (supporting TiDB Vector, Qdrant, pgvector, and others). When conflicting facts appear, Mem0 relies on the LLM to overwrite or update in place at write time. On test sets such as LongMemEval, Mem0 shows high fact-retrieval accuracy, and its out-of-the-box convenience is its biggest strength.

Letta: treating the LLM context as an operating system's memory

Letta inherits the classic idea of the MemGPT paper, likening the LLM's context window to an operating system's memory (RAM) and external storage to a hard disk, and argues that an agent should be able to self-edit its memory. It divides memory into three tiers:

  1. Core Memory: resident in the prompt context, containing the persona and key user profile; the agent can actively modify this area by calling tools.
  2. Archival Memory: an unbounded external vector store. When Core Memory cannot hold everything, the agent pages information out to the archive and pages it back into context via retrieval when needed.
  3. Recall Memory: a complete log of all conversation history.

Internally, Letta manages memory changes with a Git-like file-and-diff tracking mechanism, making it well suited to building complex agents that run for long periods and can self-evolve.

Zep / Graphiti: a bi-temporal knowledge graph

Zep and its underlying open-source engine Graphiti represent the direction of agent memory evolving toward graph theory and temporal reasoning. Zep's central judgment is that conventional vector retrieval cannot handle facts that evolve and conflict over time. Take "the user lived in Beijing last month and moved to Shanghai this month": with vector matching alone, the agent retrieves both contradictory pieces of information at once and the LLM hallucinates.

Graphiti gives every edge (relationship) in the graph a precise timestamp dimension with two layers of time: the "valid time of the fact" (Valid-From / Valid-To) and the "time the system learned the fact" (Learned-At / Expired-At). When a new contradictory fact arrives, Graphiti does not simply delete the old data; it marks the old edge's Valid-To as invalidated and creates a new edge. This bi-temporal model can answer not only a user's current state but also precisely what the system knew at a specific point in history—making it an excellent fit for finance, legal, and healthcare scenarios with hard audit and traceability requirements.

LangMem: tripartite memory fused with LangGraph

LangMem, released by the LangChain team, aims for deep, seamless coordination with LangGraph's persistence primitives (BaseStore and Checkpointer). It extends the memory abstraction into three dimensions:

  • Semantic Memory: records facts and preferences about the user or environment.
  • Episodic Memory: stores the agent's successful experiences and reasoning-chain traces (Execution Traces) from past tasks, using them as future few-shot examples.
  • Procedural Memory / System Instructions: through the Prompt Optimizer API, analyzes historical conversations and user feedback to progressively optimize and rewrite the agent's System Prompt.

Comparing the four open-source frameworks

Architecture dimensionMem0Letta (MemGPT)Zep / GraphitiLangMem (LangGraph)
Underlying data modelVector embeddings + JSON-extracted factsTiered State Blocks (Core/Archival)Bi-temporal knowledge graph (Temporal Graph)Tripartite memory (Semantic/Episodic/Procedural)
Retrieval routingSemantic similarity (Vector Search)Paging/scheduling (Core resident + Archival retrieval)Vector + hybrid text + graph traversalSemantic matching + few-shot trace recall
Conflict and invalidationLLM-based in-place overwrite at write timeAgent calls a tool to edit manually, recording Git diffsGraph edges marked with timestamp invalidation (Valid-To)Background LLM pass tidies and updates Store state
Program/behavior evolutionLimited to factual data updatesChanges persona by modifying Core MemoryExtracts higher-order Observation rulesNatively provides Prompt Optimizer to change instructions dynamically
Typical backend fitVarious vector DBs (including TiDB Vector)PostgreSQL / Git-backed StoreNeo4j / FalkorDB / Zep cloud serviceLangGraph BaseStore (Redis/Postgres)
Cloud compatibilityCloud-neutral, fully cross-cloudCloud-neutral, fully cross-cloudCloud-neutral; self-hosted and Zep CloudCloud-neutral, fits the LangGraph ecosystem

Multi-Cloud Selection: Split by Data Compliance and Task Complexity

There is no universally optimal architecture. The key is the company's existing technology stack, its data-compliance requirements, and the complexity of the business tasks. Four scenarios help split the decision.

Heavily regulated industries such as finance and healthcare: the core requirement is a strict data-compliance boundary and zero data movement. If core business data already lives in an Oracle database environment, Oracle AI Agent Memory (Oracle AI Database 23ai) is the most direct path: build the persistent memory layer directly on the company's system of record. This avoids the compliance risk of exporting sensitive data to external third-party storage and directly inherits the database's native row-level security, audit logs, and encryption.

Cloud-native applications built deeply on AWS or GCP: adopting a managed memory service (AWS Bedrock AgentCore Memory or GCP Vertex AI Memory Bank) minimizes the infrastructure burden. These options provide built-in fact extraction and automatic consolidation logic, making them suitable for fast delivery. During rollout, however, architects must explicitly bind tenant ownership at the application layer to prevent cross-tenant, unauthorized memory retrieval, and must establish memory-pruning mechanisms to control token-cost inflation from background LLM consolidation.

Multi-tenant SaaS platforms that need elastic deployment across AWS, GCP, and private clouds: building the memory layer on the distributed HTAP database TiDB Vector together with Mem0 or LangMem balances performance and cross-cloud freedom. TiDB Vector provides high-concurrency distributed SQL plus joint vector retrieval, letting a company serve many online agents with low-latency memory CRUD while using standard SQL to run real-time analytics directly over the full history of memory data.

Scenarios that depend heavily on temporal logic, such as complex supply-chain tracking, customer-state evolution, and legal-risk investigation: Zep/Graphiti's bi-temporal knowledge graph is the better fit. Conventional vector databases are prone to inducing LLM hallucinations when facts change frequently, whereas a bi-temporal graph precisely records when relationships become valid and invalid and supports point-in-time historical queries, providing reliable evolutionary traces and explainability for business scenarios that demand strict auditing.

Conclusion: From "External Patch" to "Embedded Infrastructure"

AI agent memory systems are shifting architecturally from an external bolt-on layer to deeply embedded infrastructure. Looking ahead, three core trends are worth watching.

First, the deep fusion of temporal graphs and vector retrieval. Memory-extraction mechanisms that rely solely on vector similarity are rapidly moving toward bi-temporal knowledge graphs, and future enterprise databases and managed storage engines will deeply embed temporal-graph reasoning to achieve a standardized unification of vector matching and structured graph traversal.

Second, the spread of procedural memory and autonomous optimization. Agent memory will no longer be limited to storing static semantic facts; it will store large amounts of past execution traces and successful experiences (Episodic Memory) and continuously correct the agent's System Prompt and behavioral strategy through automated tool feedback, achieving genuine self-directed evolution.

Third, the standardization of memory-control protocols and interfaces. With the evolution of open transport protocols such as the Model Context Protocol (MCP), how agent memory is stored and accessed is converging toward uniformity, which will allow memory data to flow and be shared securely and compliantly across different cloud platforms, managed services, and agent frameworks.

VIBECODING

Readable articles from the intersection of AI and real-world development.

© 2026 VibeCoding Japan, Inc. All Rights Reserved.