My CloudOps agent remembered that I preferred the production environment.

Then I opened a new conversation and asked about deployment health. It recalled production, checked the deployment status tool, and returned a fresh result.

That was useful. But it raised a more important question:

What should the agent remember, and what should it always check again?

This was a small experiment in my personal AWS lab using Amazon Bedrock AgentCore Memory. The deployment data was synthetic, and the agent did not access a real production system.

The Agent Already Had One Tool

In the previous stage, I gave the agent a read only deployment status tool:

Question -> AgentCore Harness -> AgentCore Gateway -> Deployment_status Lambda

The Lambda supported one environment: production. It returned a sample service version, instance count, alarms, health status, and a fresh observation time.

The agent could check a current result, but it couldn't remember which environment I preferred when I started a different conversation.

So I added managed AgentCore Memory with the USER_PREFERENCE strategy.

The design now had two sources:

Memory -> which environment I prefer
Tool   -> what that environment reports now

I wanted to make sure the agent didn't confuse them.

Remembering Across Sessions Took Time

I first told the agent:

My preferred environment is production.

When I asked again in the same session, it remembered immediately.

Then I started a new session for the same actor. The first attempt did not recall the preference.

At first, that looked like i have done something wrong. But i wasn't

Upon checking documentation i found that AgentCore creates long-term memories asynchronously. The conversation must be stored, processed, and turned into a preference that can be retrieved later. After I allowed time for that processing and opened another session, the agent remembered production.

It then called the deployment-status tool and returned a new health result.

That gave me the first lesson:

Long-term memory is not an immediate key-value write.

If your application needs a preference to be available instantly in a new session, you should account for that processing delay.

A Different Actor Did Not Get My Preference

Next, I used a different actor and asked for the preferred environment.

The agent did not return the original actor's production preference.

This was the result I wanted. The preference belonged to the actor, not globally to the agent.

It also clarified the difference between an actor and a session:

  • A session represents one conversation.

  • An actor represents the person or system whose memory can continue across conversations.

A session ID alone is not a durable user identity. The application still has to assign the correct actor identity.

Staging Exposed the Real Boundary

I then changed the original actor's preference from production to staging.

After the new preference was processed, another session remembered staging.

But the deployment-status tool only supported production.

This was the real test.

The agent remembered that I preferred staging, but it could not retrieve staging health. Instead of guessing, it said that the current health was unknown.

Memory: "This actor prefers staging."

Tool: "Staging is not supported."

Answer: "I cannot determine staging health."

That is exactly how I want this boundary to work.

Memory answered:

Which environment does this person care about?

The tool answered:

What does the deployment source report now?

Remembering an environment name did not make that environment healthy, supported, or current.

The Rule I Would Keep

The main lesson from this experiment is simple:

Use memory to decide what to look up. Use a current source to decide what is true.

Memory removes repetition. It should not create confidence without evidence.

What Comes Next

This stage gave the agent continuity without letting remembered context replace current evidence.

In the next stage, I use AgentCore Code Interpreter to analyze the given deployment history. That creates another boundary to test a correct looking answer is not enough unless I can inspect the generated code and the actual execution result.

References