# Smart Home Assistant Agent - Architecture & Design ## Table of Contents - [1. System Overview](#1-system-overview) - [2. High-Level Architecture](#2-high-level-architecture) - [3. Data Flow Diagrams](#3-data-flow-diagrams) - [4. Network Architecture](#4-network-architecture) - [5. Security Architecture](#5-security-architecture) - [6. Device Simulator Design](#6-device-simulator-design) - [7. Chatbot Design](#7-chatbot-design) - [8. AI Agent Design](#8-ai-agent-design) (includes [8.6 Agent Internal Data Flow](#86-agent-internal-data-flow)) - [8.7. Skill Management](#87-skill-management) - [8.7.1. Prompt Caching](#871-prompt-caching) - [8.8. Per-User Model Selection](#88-per-user-model-selection) - [8.9. Per-Login Session ID and Session Tracking](#89-per-login-session-id-and-session-tracking) - [8.10. Agent System Prompts (Text & Voice)](#810-agent-system-prompts-text--voice) - [8.11. Image Input (Vision Bypass Path)](#811-image-input-vision-bypass-path) - [8.12. AgentCore Optimization (Recommendations, Bundles, A/B Tests)](#812-agentcore-optimization-recommendations--target-based-ab-routing) - [8.13. Per-Tenant Entry Environment](#813-per-tenant-entry-environment) - [9. Infrastructure Design](#9-infrastructure-design) - [9.4. Admin Console Design](#94-admin-console-design) - [9.4.1. User Provisioning (Identity Tab)](#941-user-provisioning-identity-tab) - [9.5. Per-User Tool Permission Management](#95-per-user-tool-permission-management) - [9.5.1. SubAgent Policy](#951-subagent-policy) - [9.6. Enterprise Knowledge Base](#96-enterprise-knowledge-base) - [9.7. Voice Mode (Nova Sonic Bi-directional Streaming)](#97-voice-mode-nova-sonic-bi-directional-streaming) - [9.8. Skill ERP & AWS Agent Registry](#98-skill-erp--aws-agent-registry) - [9.9. Integration Registry & A2A Agents](#99-integration-registry--a2a-agents) - [9.10. Remote Shell Commands per Session](#910-remote-shell-commands-per-session) - [9.11. Browser Use — Live Agent Web Automation](#911-browser-use--live-agent-web-automation) - [9.12. Observability and AgentCore Online Evaluation](#912-observability-and-agentcore-online-evaluation) - [9.13. A2A Specialist Agents & Text Agent A2A Client](#913-a2a-specialist-agents--text-agent-a2a-client) - [9.14. Code Interpreter — Live Agent Code Execution](#914-code-interpreter--live-agent-code-execution) - [9.15. Agent Operations Dashboard](#915-agent-operations-dashboard) - [9.16. Simulated End Users (Test Data Generation)](#916-simulated-end-users-test-data-generation) - [9.17. Task Management & Scheduled Automations](#917-task-management--scheduled-automations) - [9.18. The Agents Page (Fleet & Per-Agent Governance)](#918-the-agents-page-fleet--per-agent-governance) - [9.19. Simulator Props: Virtual Clock, Screen and Speaker](#919-simulator-props-virtual-clock-screen-and-speaker) - [9.20. Discovery and Developer-Facing Surfaces](#920-discovery-and-developer-facing-surfaces) - [10. API Reference](#10-api-reference) - [11. MQTT Topic & Command Reference](#11-mqtt-topic--command-reference) - [12. Error Handling Strategy](#12-error-handling-strategy) - [13. Frontend Build Pipeline](#13-frontend-build-pipeline) - [14. Technology Choices and Rationale](#14-technology-choices-and-rationale) - [15. Scalability Considerations](#15-scalability-considerations) --- ## 1. System Overview The Smart Home Assistant Agent is a full-stack application that demonstrates AI-driven smart home device control on AWS, with the **Agent Harness control plane** — governance of prompts, skills, tool authorisation, A/B testing, observability, evaluation and identity — as its point. Subsystems: | Subsystem | Technology | Purpose | |-----------|-----------|---------| | Device Simulator | React + TypeScript + Cloudscape + MQTT + Cognito | Per-user authenticated visual simulation of the device fleet declared in `shared/device-catalog.json` (12 devices across 3 rooms); each user sees only their own scope via `smarthome///...` topics | | Chatbot | React + TypeScript + Cloudscape + HTTP POST | Natural-language interface to the AI agent. Fully Cloudscape-native UI (including bubbles/input), renders agent replies as markdown, supports image attachments (≤3 images, ≤20 MB each) that route to a vision model | | AI Agent (text) | Strands Agent on AgentCore Runtime `smarthome` (Claude Sonnet 4.6 `us.anthropic.claude-sonnet-4-6` default; per-user Bedrock model override, e.g. Kimi K2.5) | Text chat via `POST /invocations`; wraps per-user MCP tools (control_device, discover_devices, query_knowledge_base) to inject the validated `user_id`; image turns bypass the text model and return the vision model's caption | | AI Agent (voice) | Strands BidiAgent on AgentCore Runtime `smarthomevoice` (Nova Sonic) | Bi-directional voice streaming via `/ws`; finalized transcripts persisted to the same AgentCore Memory the text agent uses | | AI Agent (vision) | Claude Haiku 4.5 default (per-user multimodal model override) via Bedrock Converse | Captions uploaded images; captions injected as prior assistant messages so the text agent can answer follow-up questions | | AI Agent (browser) | `browser-use` + AgentCore Browser Tool (`aws.browser.v1`) driven by the text agent's `browse_web` Strands tool | Live web automation when the user asks something that needs a real site (product search, news, Wikipedia). Chatbot renders the live DCV stream + a Take/Release Control button; per-step screenshots land in the agent session's `/mnt/workspace//browser/` (see §9.11). | | AI Agent (code) | `code-interpreter` + AgentCore Code Interpreter (`aws.codeinterpreter.v1`) driven by the text agent's `execute_python` Strands tool | Live Python execution for data analysis, optimization, simulation, and charting over smart-home telemetry. Chatbot's right-side panel auto-opens a "CodeInterpreter" tab rendering each block's code + streamed stdout/stderr + inline matplotlib charts; charts land in `/mnt/workspace//code/` (see §9.14). | | Tool Access | AgentCore Gateway (MCP Server) + Lambda + curated Strands built-ins | Device discovery, command routing, KB query, and device control via MCP. Built-in Strands/AgentCore tools (`http_request`, `file_write`, etc.) also surfaced for admin per-user policy and for reference skills. | | AI Agents (specialists) | 8 independent AgentCore Runtimes reached over the A2A protocol | Domain specialists the orchestrator delegates to: device control, lighting effects, knowledge QA, task management (stored scenes and routines), security, energy, appliance maintenance, live scene sync. Each carries the caller's verified identity to the same Gateway, so Cedar evaluates the real end user (see §9.13). | | Task management | task-management A2A agent + `smarthome-scenarios` table + EventBridge Scheduler + `smarthome-scenario-runner` Lambda | Turns a described routine into a stored scene (trigger + device actions) and executes it on time **as the owner, through the Gateway**, so a scheduled command is authorised exactly like a hand-typed one. Five trigger kinds: clock time in the owner's timezone, sunrise/sunset from their coordinates, device state, sensor threshold, and one-tap manual (see §9.17). | | Live scene sync | scene-sync A2A agent + the simulator's screen/music props | Drives a room in real time — backlight following the picture, lights on the beat — and waits on the reported Bluetooth link instead of assuming a pairing succeeded. Holds device tools and no table, the mirror of task-management (see §9.17). | | Admin Console | React + TypeScript + Cloudscape + REST API | Agent Harness Control Center with AWS-Console-style left-nav: Discover (Overview, **Agents**, Integration Registry), Build (Models, Skills, Prompt, Tool Policy, Memories, Knowledge Base, Identity), Deploy (Instance Type, Sessions), Assess (Agent Guardrails, Observability, Evaluations, Optimization). Supports light/dark themes. | | Skill ERP | React + TypeScript + Cloudscape + REST API | End-user skill + A2A agent publishing: authors SKILL.md and A2A records, publishes to AWS Agent Registry for curator approval | | Enterprise Knowledge Base | Bedrock KB + **S3 Vectors** + S3 | RAG retrieval with per-user document isolation via S3 prefix + metadata filtering. Vector store is the pay-per-vector S3 Vectors service (no fixed monthly floor). | | Infrastructure | AWS CDK (TypeScript) | One-click deployment of all resources | --- ## 2. High-Level Architecture ``` +--------------------------------------------------------------------------------------------+ | AWS Cloud | | | | +--------------+ +--------------+ +--------------+ +---------------------------+ | | | CloudFront | | CloudFront | | CloudFront | | Cognito | | | | (Device Sim) | | (Chatbot) | | (Admin) | | +--------+ +--------+ | | | +------+-------+ +------+-------+ +------+-------+ | | User | |Identity| | | | | | | | | Pool | | Pool | | | | +------v-------+ +------v-------+ +------v-------+ | +---+----+ +---+----+ | | | | S3 Bucket | | S3 Bucket | | S3 Bucket | +------+----------+--------+ | | +--------------+ +--------------+ +------+-------+ | | | | | | | | | | +-----------+ MQTT/WSS | Bearer JWT | Bearer JWT | | | | | Browser +--------+ | | | | | | | (Device | | +-----v----------v---+ | | | | | Sim App) | | | AgentCore Runtime | +----------v-------+ | | | +-----------+ | | (Strands Agent) | | API Gateway | | | | | | +-----+---------------+ | (Admin API) | | | | Cognito | | | +----+-------------+ | | | Identity | | | MCP Client | | | | Pool | | +-----v-----------------+ | | | | (SigV4) | | | AgentCore Gateway | | | | | v | +-----+-----------------+ +--v-------------+ | | | +------------------+ | | | admin-api | | | | | AWS IoT Core | | | Lambda Targets | Lambda | | | | | MQTT Broker |<+--------+ +--+-------------+ | | | +------------------+ | +-----v-----------------+ | | | | | | iot-control Lambda | +--v-------------+ | | | | | Validates & publishes | | DynamoDB | | | | +->| MQTT | | (skills table) |<-+ Agent reads | | | +------------------------+ +----------------+ | | | +------------------------+ +----------------+ | | | | iot-discovery Lambda | | S3 Bucket | | | | | Returns device list | | (skill files) |<-- Admin Lambda | | | +------------------------+ +----------------+ (presigned URLs) | +--------------------------------------------------------------------------------------------+ ``` ### Component Interaction Matrix ``` Cognito Cognito IoT AgentCore AgentCore IoT Control IoT Discovery Admin DynamoDB S3 Skill UserPool Identity Core Runtime Gateway Lambda Lambda Lambda Skills Files Pool Device Simulator R R/W Chatbot App R R/W Admin Console R R/W AgentCore Runtime V I (MCP) R AgentCore Gateway V (self) I I IoT Control Lambda W IoT Discovery Lambda (self) Admin Lambda (self) R/W R/W R = Read/Subscribe W = Write/Publish V = Validate JWT I = Invoke ``` --- ## 3. Data Flow Diagrams ### Flow 1: Device Command ``` User (Chatbot, logged in — idToken carries Cognito sub) | | "Turn on the LED matrix to rainbow mode" v AgentCore Runtime (HTTP POST /invocations) --> Strands Agent (Claude Sonnet 4.6 default) | | Agent wraps control_device / discover_devices so the | JWT-validated `sub` is injected as `user_id` before | the MCP call — LLM cannot forge this argument. | | MCP Client call (with user_id) v AgentCore Gateway (MCP Server) | | Lambda Target v iot-control Lambda | | Refuses requests without user_id. | iot-data:Publish scoped to the caller's sub. v AWS IoT Core Topic: smarthome//living-led-1/command | | MQTT over WebSocket (SigV4) v Device Simulator (Browser, same signed-in user) the catalog-rendered LED matrix component receives: {"action":"setMode","mode":"rainbow"} ``` ### Flow 2: Device Discovery ``` User (Chatbot) | | "Turn on all my devices" v AgentCore Runtime --> Strands Agent (activates all-devices-on skill) | | 1. MCP call: discover_devices() v AgentCore Gateway | | Lambda Target v iot-discovery Lambda | | Returns the device catalog (deviceId, room, capabilities, actions) v Agent receives device list | | 2. For each device (sequentially, 5s apart): | MCP call: control_device(device_id, setPower command) v iot-control Lambda --> IoT Core --> Device Simulator ``` ### Flow 3: User Authentication ``` User (Browser) | | email + password v Chatbot LoginPage | | amazon-cognito-identity-js v Cognito User Pool | | returns: idToken, accessToken, refreshToken v ChatInterface | | HTTP POST to AgentCore Runtime /invocations | Authorization: Bearer {idToken} v AgentCore Runtime | | JWT validation (Cognito User Pool) v Strands Agent processes request ``` ### Flow 4: Device Simulator MQTT Connection ``` Device Simulator (Browser) | | 1. Cognito User Pool sign-in -> idToken (JWT with `sub`) v LoginPage (Cloudscape) | | 2. Attach `smarthome-device-sim-client` IoT policy to this identity | (workaround: IoT refuses SigV4 MQTT over WS without an attached IoT policy) v Cognito Identity Pool (authenticated, federated via User Pool login) | | 3. Temporary AWS credentials (authenticated role) v MqttClient.ts | | 4. MQTT5 over WebSocket with SigV4 | ClientId: "device-sim-{random}" v AWS IoT Core | | 5. Subscribe (scoped to the signed-in user's own sub), one topic per | catalog device (shared/device-catalog.json): | smarthome///command | e.g. smarthome//living-led-1/command | smarthome//kitchen-oven-1/command v Each device component receives commands and updates its UI state ``` --- ## 4. Network Architecture ``` Internet | +---> CloudFront (Device Simulator) ---> S3 Bucket (static assets) | | | +---> /config.js (runtime config from S3) | +---> CloudFront (Chatbot) ---> S3 Bucket (static assets) | | | +---> /config.js (runtime config from S3) | +---> HTTPS: AgentCore Optimization Gateway (chatbot text path — see §8.12) | https://{opt-gw-id}.gateway.bedrock-agentcore.{region}.amazonaws.com | /smarthome-control/invocations | | | +---> Routes by gatewayFilter + active A/B test to runtime endpoint | | "control" or "treatment" on the smarthome runtime. | +---> Auth: AWS SigV4 (Cognito Identity Pool authenticated role, | with bedrock-agentcore:InvokeGateway permission). | +---> WSS: AgentCore Runtime (voice WebSocket — direct, gateway can't proxy WS) | wss://bedrock-agentcore.{region}.amazonaws.com/runtimes/{voiceArn}/ws | | | +---> Strands BidiAgent (Nova Sonic) on `smarthomevoice` runtime | +---> Auth: AWS SigV4 presigned URL | +---> HTTPS: AgentCore Runtime (warmup ping, /ping health check) | https://bedrock-agentcore.{region}.amazonaws.com/runtimes/{textArn}/invocations | | | +---> __warmup__ POST direct to text runtime (heats microVM ahead | of first chat turn; not used for real chat traffic anymore). | +---> CloudFront (Admin Console) ---> S3 Bucket (static assets) | | | +---> /config.js (adminApiUrl, Cognito IDs) | +---> HTTPS: Admin API (API Gateway) | https://{api-id}.execute-api.{region}.amazonaws.com/prod/skills | | | +---> Lambda Target: admin-api Lambda | +---> Auth: Cognito User Pool Authorizer (admin group required) | +---> DynamoDB: smarthome-skills table (skill CRUD, all spec fields) | +---> S3: smarthome-skill-files bucket (file management via presigned URLs) | +---> HTTPS: AgentCore Gateway (MCP Server, internal to agent) | https://{gateway-id}.gateway.bedrock-agentcore.{region}.amazonaws.com/mcp | | | +---> Lambda Target: iot-control Lambda (control_device tool) | +---> Lambda Target: iot-discovery Lambda (discover_devices tool) | +---> Auth: CUSTOM_JWT (Cognito, validates user JWT for per-user Cedar policies) | +---> Policy Engine: Cedar per-user tool access control (ENFORCE mode) | +---> WSS: IoT Core (iot-endpoint.iot.region.amazonaws.com) | +---> MQTT5 over WebSocket (SigV4 auth via Cognito Identity Pool) +---> Topics: smarthome/{device_type}/command ``` --- ## 5. Security Architecture ``` +------------------------------------------------------------------+ | Authentication Layers | +------------------------------------------------------------------+ | | | Layer 1: Cognito User Pool | | +------------------------------------------------------------+ | | | - Email/password authentication | | | | - Self-service sign-up with email verification | | | | - Issues JWT tokens (id, access, refresh) | | | | - Used by: Chatbot app, AgentCore Runtime, AgentCore Gateway| | | +------------------------------------------------------------+ | | | | Layer 2: Cognito Identity Pool | | +------------------------------------------------------------+ | | | - Federation with User Pool | | | | - Issues temporary AWS credentials via STS | | | | - Unauthenticated access allowed (for device sim) | | | | - Scoped IAM role: iot:Connect/Subscribe/Receive/Publish | | | | - Used by: Device simulator MQTT connection | | | +------------------------------------------------------------+ | | | | Layer 3: AgentCore Runtime AWS_IAM (SigV4) Authorization | | +------------------------------------------------------------+ | | | - Browser signs with Cognito Identity Pool authenticated | | | | role credentials (service=bedrock-agentcore) | | | | - SigV4 headers on POST /invocations | | | | - SigV4 presigned URL (X-Amz-* in query) on WebSocket /ws | | | | - Protects: /invocations, /ping, /ws endpoints | | | +------------------------------------------------------------+ | | | | Layer 4: AgentCore Gateway (CUSTOM_JWT + Policy Engine) | | +------------------------------------------------------------+ | | | - Auth: CUSTOM_JWT (same Cognito User Pool) | | | | - Chatbot ships idToken in custom header | | | | X-Amzn-Bedrock-AgentCore-Runtime-Custom-AuthToken | | | | (allowlisted via requestHeaderAllowlist) | | | | - Agent forwards it as Bearer to gateway MCP client | | | | - Policy Engine: Cedar per-user permit policies (ENFORCE) | | | | principal.id from JWT sub claim for per-user control | | | +------------------------------------------------------------+ | | | | Layer 5: Admin API (Cognito + Group Check) | | +------------------------------------------------------------+ | | | - API Gateway Cognito User Pools Authorizer (validates JWT) | | | | - Lambda checks cognito:groups claim for "admin" membership | | | | - Protects: /skills CRUD endpoints | | | | - Non-admin users receive 403 Forbidden | | | +------------------------------------------------------------+ | | | +------------------------------------------------------------------+ ``` ### IAM Permissions (Least Privilege) | Principal | Permissions | Scope | |-----------|------------|-------| | Cognito Unauth Role | iot:Connect, Subscribe, Receive, Publish | `*` (IoT Core) | | Cognito Auth Role | iot:Connect, Subscribe, Receive, Publish | `*` (IoT Core) | | iot-control Lambda | iot:Publish | `arn:...:topic/smarthome/*` | | iot-discovery Lambda | (none) | Returns the device catalog (`shared/device-catalog.json`) | | admin-api Lambda | dynamodb:* | `arn:...:table/smarthome-skills` | | admin-api Lambda | dynamodb:Scan, GetItem, UpdateItem, DeleteItem | `arn:...:table/smarthome-runtime-sessions` | | admin-api Lambda | dynamodb:PutItem, Scan | `arn:...:table/smarthome-feedback` (write a vote, aggregate for the dashboard) | | admin-api Lambda | s3:GetObject, PutObject, DeleteObject, ListBucket | `arn:...:smarthome-skill-files-*` | | admin-api Lambda | cognito-idp:ListUsers, AdminListGroupsForUser | Cognito User Pool | | admin-api Lambda | bedrock-agentcore:ListActors, ListMemoryRecords | `*` (AgentCore Memory) | | admin-api Lambda | bedrock-agentcore:Create/Get/Update/Delete Policy* | `*` (policy engine + gateway management) | | admin-api Lambda | iam:PassRole, iam:PutRolePolicy | AgentCore roles | | AgentCore Runtime Role | bedrock:InvokeModel (+ WithResponseStream / WithBidirectionalStream) | `*` (default Claude Sonnet 4.6, per-user overrides, vision and Nova Sonic models) | | AgentCore Runtime Role | dynamodb:Query, GetItem, Scan | `arn:...:table/smarthome-skills` | | AgentCore Runtime Role | dynamodb:PutItem, UpdateItem | `arn:...:table/smarthome-runtime-sessions` (records each per-login session) | --- ## 6. Device Simulator Design ### 6.1 MQTT Connection Strategy The device simulator runs entirely in the browser. Users sign in with the shared Cognito User Pool (same one the chatbot uses) and the resulting idToken federates authenticated Identity Pool credentials for SigV4 MQTT. ``` Browser | +-> LoginPage (Cognito User Pool: signIn / signUp / confirm) | -> idToken (JWT with `sub` = user UUID) | +-> AttachPolicy(smarthome-device-sim-client) to the caller's Cognito identity | (workaround: IoT refuses SigV4 WS connects without an attached IoT policy, | even when IAM already allows iot:Connect) | +-> fromCognitoIdentityPool({ logins: { 'cognito-idp...': idToken } }) | -> authenticated temporary AWS credentials | +-> Mqtt5Client (aws-iot-device-sdk-v2) |-> WebSocket SigV4 auth |-> clientId: "device-sim-{random}" |-> keepAlive: 30s |-> auto-reconnect with re-subscribe, refresh every 45min |-> subscribes to: smarthome///command (one per catalog device) ``` **Note on the browser SDK:** The aws-iot-device-sdk-v2 has different APIs for Node.js and browser bundles. The browser build exports `auth.StaticCredentialProvider` (not `auth.AwsCredentialsProvider.newStatic()`). TypeScript uses `(auth as any).StaticCredentialProvider` since the Node.js type definitions don't include the browser API. **Key design decisions:** - **Authenticated access required**: The device simulator gates on Cognito sign-in. The unauthenticated Identity Pool role exists as a CDK remnant but the new client never uses it. - **Per-user MQTT topics**: All subscriptions and publishes use `smarthome///command` (and `/state` for reported state). Each user's browser sees only their own simulated devices. - **IoT AttachPolicy at login**: `ensureIotPolicyAttached()` runs on sign-in before the MQTT client starts. It attaches the `smarthome-device-sim-client` IoT policy (CDK-managed) to the user's Cognito identity ID. This is AWS's documented workaround for authenticated Cognito users getting `Connection refused: Not authorized` when trying to use MQTT over WebSockets. - **Singleton MQTT client**: All device components share one `MqttClient` instance to avoid multiple WebSocket connections. - **All devices off by default**: Every device starts in the powered-off state when the page loads. - **Auto-power-on**: Setting a mode, speed, temperature, or color via MQTT automatically powers on the device (declared per action as `implies` in `shared/device-catalog.json`, e.g. `setSpeed` implies `power: true`), so the agent doesn't need to send a separate `setPower` command. - **Layout**: one card per catalog device in a packed dashboard grid (`dashboard-grid`), labelled with its room. #### Theming: never reference Cloudscape's own CSS variables from app CSS Cloudscape emits component-scoped custom properties with a **build-time hash suffix**, so `var(--color-background-container-content, #fff)` in app CSS never resolves and the **fallback silently wins**. Both the simulator and the chatbot therefore sample the resolved light/dark values and push them onto `document.body` as stable aliases (`--sim-*`, `--chat-*`) in `src/theme/applyTheme.ts`; app CSS must use only those. This failed exactly as designed to fail. The Virtual Clock and Screen-and-speaker panels still referenced the raw `--color-*` names, so both stayed **white in dark mode** while every device card beside them flipped correctly — and because a fallback is a legitimate-looking colour, it read as a deliberate light surface rather than a broken reference. Nothing warned: an unresolvable `var()` is not an error in CSS, it is a fallback. The lesson generalises past this repo: **a defaulted lookup and a correct lookup are indistinguishable in the output.** The fix is mechanical (use the aliases), and the guard is now a test rather than a habit: `shared/tests/test_theme_tokens.py` fails on any `var(--color-` in an app's CSS, requires all three apps to publish aliases for both modes, and **measures the published palettes for WCAG AA contrast**. **The second failure mode is a hardcoded colour that only suits one mode**, and it is the one that produced a bug report. The admin console's `App.css` was written dark-only and published no aliases at all, so ~65 rules carried literals like `#e0e0e0` and `#8888aa`. Correct on a dark surface; on white, `.perm-tool-name` measured **1.32:1** (AA wants 4.5:1). The Tool Policy permission list therefore looked greyed out and was reported as "per-user tool permissions no longer work" — while every one of the 17 checkboxes was enabled and interactive and `/tools` returned all 17. A contrast failure is indistinguishable from a disabled control, which is why this is in the silent-success family rather than a cosmetic issue. Contrast has to be **measured, not eyeballed**: the first two values chosen for the dim tier came out at 3.34:1 and 4.49:1, both below the floor and both fine to the eye. Two things the test deliberately does not flag, because they are correct: the ANSI palette and the remote shell's foregrounds (they render on hardcoded dark terminal panes in both themes), and white text on filled buttons. Distinguishing those from genuine bugs meant checking placement **in the components** — `ShellModal.tsx` renders in the modal body, `AnsiOutput.tsx` renders in the dark pane — because an alias sweep driven by CSS order alone produced dark-on-dark text in three rules. #### Security model & known limitation Per-user isolation has two layers: - **Agent → Lambda path (strict)**: the `iot-control` Lambda publishes only on `smarthome//...`. The `user_id` comes from the agent's wrapper, which decodes it from the runtime-validated idToken — the LLM cannot forge it. - **Browser → IoT Core (best-effort)**: the authenticated Cognito IAM role currently allows IoT operations on `*` because mapping Identity Pool identity IDs back to User Pool subs is nontrivial. A determined user could subscribe to another user's topics with raw MQTT; closing that gap would require an IoT Core Policy keyed on a Cognito custom claim and is deferred as future hardening. ### 6.2 Device Components The fleet is declared once in `shared/device-catalog.json` — 12 devices across 3 rooms (living room, bedroom, kitchen): a light strip, a bedroom light, a fan, a plug, a temperature/humidity sensor, an LED matrix, a rice cooker, a humidifier, an air purifier, an ice maker, a TV backlight and an oven. Each entry carries a `deviceId` of the form `{room}-{type}-{seq}` (e.g. `living-led-1`), which is both the MQTT topic segment and the agent's handle for the device, plus the `capabilities` (typed, bounded) and `actions` (wire action name → capability it writes, with `implies` side effects such as `setSpeed` implying `power: true`). The same file drives three consumers: the simulator renders a card per entry, `iot-control` validates (and clamps) commands against it, and `iot-discovery` returns it to the agent. Every component shares one state hook, `useDeviceState(device, userSub)` (`device-simulator/src/state/useDeviceState.ts`), which subscribes to `smarthome/{userSub}/{deviceId}/command`, applies the action through the catalog's `actions` map, and reports the resulting state back on `smarthome/{userSub}/{deviceId}/state`. A UI click and an agent command take the same path out. `App.tsx` chooses a component by shape rather than by device: | Shape | Component | Devices | |-------|-----------|---------| | Read-only capabilities | `SensorDevice.tsx` | sensor | | Has `color` or `segments` | `LightDevice.tsx` (shared effect engine in `devices/effects.ts`, `requestAnimationFrame` loop; strip = 30 segments, matrix = 16x16 = 256, bulb = 1) | light strip, bedroom light, LED matrix, TV backlight | | `fan` | `Fan.tsx` | fan | | `oven` | `Oven.tsx` | oven | | `rice_cooker` | `RiceCooker.tsx` | rice cooker | | Anything else | `GenericDevice.tsx` (toggle / level meter / button row generated from capabilities) | plug, humidifier, air purifier, ice maker | `DeviceCard.tsx` is the shared card shell (title, room, status dot). The TV backlight additionally drives the `MediaSync.tsx` screen-and-speaker panel (§9.19). #### Fan (Fan.tsx) CSS-animated spinning fan. Speed is an integer `0-8` from the catalog (`capabilities.speed.max`); the spin duration is computed from the ratio `speed / maxSpeed` rather than a per-step table, so the visual and the Lambda's clamp agree. Oscillation adds a secondary `fan-osc` animation (4s ease-in-out oscillation). #### Rice Cooker (RiceCooker.tsx) | State | Behavior | |-------|----------| | idle | Display shows "Ready", temperature cools to 25C | | cooking | Timer counts down, temperature rises to target, steam animation plays | | keep_warm | Maintains 65C after cooking completes | | done | Shows "DONE" on display | `cooking`, `mode` and `keep_warm` are catalog-backed; the four-way status above is local and derived. #### Oven (Oven.tsx) Visual elements: - Control panel with knobs and digital display - Glass window with visible heating elements (top and bottom) - Gradient glow that rises proportionally to temperature Temperature simulation: heats at 5% of remaining delta per tick (500ms), with +/-1F fluctuation at target. Auto-transitions from Preheat to Bake when within 5F of target. ### 6.3 Runtime Configuration Configuration is injected at deploy time via a `config.js` file served from S3: ```javascript // config.js (written by CDK custom resource AND re-written by setup-agentcore.py) window.__CONFIG__ = { iotEndpoint: "a1b2c3d4e5f6g7-ats.iot.us-east-1.amazonaws.com", region: "us-east-1", cognitoIdentityPoolId: "us-east-1:xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx", cognitoUserPoolId: "us-east-1_XXXXXXXXX", cognitoClientId: "abcdef1234567890" }; ``` This approach avoids baking environment-specific values into the webpack bundle, enabling the same build to work across environments. The CDK stack's `BucketDeployment` for the device-sim bundle uses `prune: false` so the `AwsCustomResource`-written `config.js` survives every deploy. `setup-agentcore.py` rewrites the same file in step 6 (to refresh IoT endpoint / identity pool values); both writers must emit the full set of Cognito fields — missing `cognitoUserPoolId`/`cognitoClientId` breaks sign-in with "Both UserPoolId and ClientId are required". --- ## 7. Chatbot Design ### 7.1 Authentication Flow ``` +-------------+ | LoginPage | +------+------+ | +-------------+-------------+ | | Sign In Sign Up | | +-----v-----+ +-------v-------+ | Cognito | | Cognito | | authUser | | signUp | +-----+------+ +-------+-------+ | | | +------v-------+ | | Confirm Code | | | (email) | | +------+-------+ | | +-------------+-------------+ | +------v------+ | AuthTokens | | stored in | | localStorage| +------+------+ | +------v------+ |ChatInterface| +-------------+ ``` **Session management:** - Tokens are managed by `amazon-cognito-identity-js` (stored automatically in localStorage) - `getCurrentSession()` checks for valid session on app mount - Tokens auto-refresh via Cognito SDK's built-in refresh flow **Login prefill via query param:** `LoginPage` reads `?username=` (or `?email=`) from `window.location.search` on mount and prefills the username field. The Admin Console's Tool Access tab uses this to launch pre-filled chatbot demos for any Cognito user, so administrators only need to enter the password. ### 7.2 Message Architecture The chatbot communicates with the AgentCore Runtime via HTTP POST. Which entry it uses is chosen per user by the tenant's entry environment (`getTenantMode(userId)`, see §8.13): ``` default Browser --HTTP POST (SSE)--> smarthome runtime ----------------> Strands Agent (Claude Sonnet 4.6 default) ab-targets Browser --HTTP POST--> Optimization Gateway --[gateway target]--> smarthome runtime (smarthome-control or smarthome-treatment) ab-bundles Browser --HTTP POST--> bundles runtime (separate runtime ARN) \-> Vision model when images present (all modes) ``` **HTTP POST Invocation:** - Endpoint (`default`, and the fallback whenever the selected mode's config field is empty): direct runtime invocation at `https://bedrock-agentcore.{region}.amazonaws.com/runtimes/{encodedArn}/invocations`. Text turns set `"stream": true` in the body and the runtime replies with SSE progress events instead of a JSON body (§9.20). - Endpoint (`ab-targets` only): the **dedicated optimization gateway** (target-based A/B routing — see §8.12), `https://{opt-gw-id}.gateway.bedrock-agentcore.{region}.amazonaws.com/smarthome-control/invocations`; the gateway routes to the `control` or `treatment` runtime endpoint by `gatewayFilter` + active A/B test config. The two A/B modes keep the JSON (non-streaming) path, because neither is guaranteed to pass a streaming body through unbuffered. - Authentication: AWS SigV4 (service `bedrock-agentcore`), signed in the browser with temporary credentials from the Cognito Identity Pool authenticated role. For `ab-targets` the role needs `bedrock-agentcore:InvokeGateway` on the optimization gateway ARN (granted by `setup-agentcore.py`). - Session ID: `X-Amzn-Bedrock-AgentCore-Runtime-Session-Id: user-session-{cognito-sub}-{Date.now()}` — also used by the optimization gateway for sticky variant assignment during A/B tests. - User ID: Passed in the POST body as `userId`. - Gateway idToken passthrough: `X-Amzn-Bedrock-AgentCore-Runtime-Custom-AuthToken: {cognito_idToken}` header, forwarded by the runtime to the agent and re-wrapped as `Bearer` on the **tools** gateway MCP client (per-user Cedar evaluation). - CORS: Fully supported. **Voice path is unchanged** — `wss://bedrock-agentcore.{region}.amazonaws.com/runtimes/{voiceArn}/ws` direct to the voice runtime. AgentCore Gateway proxies HTTP only, not WebSocket. **WebSocket (voice mode):** See §9.7 for the full flow. Same host, path `/ws`, signed via SigV4 presigned URL; session-id + AuthToken travel as signed query parameters because browsers can't set custom headers on a WebSocket handshake. #### HTTP Message Protocol **Client -> Server (POST) — text only:** ```json {"prompt": "Turn on the LED to rainbow mode", "userId": "user@example.com"} ``` **Client -> Server (POST) — with image attachments:** ```json { "prompt": "describe this dashboard", "userId": "user@example.com", "images": [ {"mediaType": "image/png", "data": ""}, {"mediaType": "image/jpeg", "data": ""} ] } ``` - `images` is optional. When present it must be a list of ≤3 objects; oversize or wrong-type items are rejected client-side before send (≤20 MB raw per image; `image/png|jpeg|webp|gif`). - When `images` is present the payload is handled by the agent's vision bypass path (see §8.11); the text model is not called. **Server -> Client (Response):** ```json {"response": "I'll set the LED matrix to rainbow mode. The command has been sent!", "status": "success"} ``` The response shape is the same for text and image turns; for image turns the body is the vision model's description (plus any `Note: …` warnings from partial failures). #### UI Pattern The `ChatInterface` has a paperclip button next to the send button that opens a multi-select file picker (up to 3 images, ≤20 MB each, `image/png|jpeg|webp|gif`). Selected images render as 48×48 thumbnails above the textarea, each with a × to remove. On send, files are base64-encoded in parallel and included in the POST body's `images` array; the user bubble also renders the thumbnails alongside the text. Typing indicator and response rendering are identical to text-only turns. To the right of the chat column the `BrowserPanel` surfaces live agent browser sessions. It defaults to a 40-px-wide vertical rail ("Browser" + "Files" labels); clicking either expands the panel to 720 px on that tab. While the agent is typing the chatbot polls `/sessions?action=browser-active` every 1.5 s; once a session is running it renders the DCV live-view stream, a Take/Release Control toggle, and a Files tab that walks the text-agent runtime's `/mnt/workspace//` via direct `InvokeAgentRuntimeCommand` SDK calls. A maximize button in the panel header flips to `flex: 1` so wide pages (Amazon, Wikipedia) render without horizontal clipping. See §9.11 for the full design. ### 7.3 Session Persistence - **Auth tokens**: Managed by Cognito SDK in `localStorage`. Auto-refreshed. - **Chat history**: Held in React state (not persisted). Refreshing the page clears chat history. - **Session ID**: Fixed per user, derived from `cognito:sub` UUID. Sent via `X-Amzn-Bedrock-AgentCore-Runtime-Session-Id` header. Same user always gets the same runtime session. - **User ID**: User's email extracted from JWT and sent in the POST body as `userId`. Used for per-user skill loading, model selection, and session tracking. - **AgentCore Memory**: Provides conversation persistence across sessions via semantic, summary, user preference, and episodic extraction strategies. --- ## 8. AI Agent Design ### 8.1 Strands Agent on AgentCore Runtime The agent is a Strands Agent deployed to Amazon Bedrock AgentCore Runtime via CodeZip (managed Python runtime). | Property | Value | |----------|-------| | Foundation Model | `us.anthropic.claude-sonnet-4-6` (Claude Sonnet 4.6 via its cross-region inference profile; `MODEL_ID` env, default in `agent/agent.py` and `scripts/setup-agentcore.py` `DEFAULT_MODEL_ID`). Other Bedrock models, e.g. `moonshotai.kimi-k2.5`, are per-user overrides (§8.8) | | Runtime Framework | Strands Agents SDK + `bedrock-agentcore` Python package | | Packaging | CodeZip (Python 3.14, `PYTHON_3_14` managed runtime — no Docker) | | Endpoints | `/invocations` (POST), `/ping` (health) on port 8080 | | App Framework | `BedrockAgentCoreApp` from `bedrock-agentcore` | | Memory | AgentCore Memory with semantic, summary, user preference, and episodic strategies | **System instruction:** > You are a smart home assistant. [...] Device control — turn devices on/off, set brightness, colour, mode, speed or temperature. Call discover_devices for the fleet and its valid parameters; never recite devices from memory. [...] Abridged; the full hardcoded text is `SYSTEM_PROMPT` in `agent/agent.py` (capabilities list, scope rules, A2A routing rules — see §8.10). The prompt deliberately names no devices: the fleet comes from `discover_devices`, which returns `shared/device-catalog.json`. ### 8.2 AgentCore Memory The agent uses AgentCore Memory for short-term conversation persistence and long-term knowledge extraction. Memory is created and deployed via the `agentcore` CLI as a first-class project resource (`agentcore add memory`), which manages its lifecycle through the same CloudFormation stack as the runtime and gateway. **All four built-in strategies:** | Strategy | Type | Namespace | |----------|------|-----------| | **Semantic** | `SEMANTIC` | `/users/{actorId}/facts` | | **Summarization** | `SUMMARIZATION` | `/summaries/{actorId}/{sessionId}` | | **User Preference** | `USER_PREFERENCE` | `/users/{actorId}/preferences` | | **Episodic** | `EPISODIC` | `/strategy/{memoryStrategyId}/actor/{actorId}/` | Episodic is the ordered account of an interaction ("started dinner, set the cooker, then turned everything off"), which summarization's prose and semantic's standalone facts both discard. Two things about it differ from the other three, both documented AWS behaviour and both discovered the hard way: - **Its namespace is strategy-scoped, not user-scoped.** Episodes live under `/strategy/{memoryStrategyId}/...`; the actor-level variant above keeps one user's episodes out of another's. A `/users/{actorId}/episodes` namespace is **accepted** by the API and then never populated — measured: from the same six events, semantic and user-preference produced records in ~50s while the mis-namespaced episodic produced none in six minutes. Because the id is minted with the strategy, it is resolved at deploy time into `MEMORY_STRATEGY_EPISODIC_ID`, and `agent/memory/session.py` **omits** the entry when that is unset rather than guessing a path that would silently retrieve nothing forever. - **Records appear only when an episode is judged complete.** Per the docs, "if an episode is not complete, it will take longer to generate because the system waits to see if the conversation is continued." An empty episodic namespace mid-session is therefore expected and is *not* evidence of a misconfiguration — which is why `agent/tests/test_memory_namespaces.py` asserts the wiring statically instead of querying the service. Adding EPISODIC to an **existing** memory needs `UpdateMemory(memoryStrategies={"addMemoryStrategies": [...]})`; the API rejects an episodic namespace that is not at or under the strategy's reflection namespace ("must be the same as or a hierarchical prefix of"), so pass `reflectionConfiguration.namespaces` with the same value. Recreating the memory instead would discard every stored record. **CLI-managed lifecycle:** ```bash # Memory is added to the agentcore project alongside gateway and runtime agentcore add memory --name SmartHomeMemory \ --strategies SEMANTIC,SUMMARIZATION,USER_PREFERENCE,EPISODIC agentcore deploy -y --verbose # CLI auto-sets MEMORY_SMARTHOMEMEMORY_ID env var on the runtime ``` **Agent integration** (`agent/memory/session.py`): ```python from bedrock_agentcore.memory.integrations.strands.config import AgentCoreMemoryConfig, RetrievalConfig from bedrock_agentcore.memory.integrations.strands.session_manager import AgentCoreMemorySessionManager MEMORY_ID = os.getenv("MEMORY_SMARTHOMEMEMORY_ID", "") # auto-set by agentcore CLI def get_memory_session_manager(session_id, actor_id): actor_id = _sanitize_actor_id(actor_id) # replace @/. with _ for Memory API retrieval_config = { f"/users/{actor_id}/facts": RetrievalConfig(top_k=3, relevance_score=0.5), f"/summaries/{actor_id}/{session_id}": RetrievalConfig(top_k=3, relevance_score=0.5), f"/users/{actor_id}/preferences": RetrievalConfig(top_k=3, relevance_score=0.5), } # Strategy-scoped, and skipped rather than guessed when the id is unknown. if EPISODIC_STRATEGY_ID: retrieval_config[f"/strategy/{EPISODIC_STRATEGY_ID}/actor/{actor_id}/"] = \ RetrievalConfig(top_k=3, relevance_score=0.5) return AgentCoreMemorySessionManager( AgentCoreMemoryConfig(memory_id=MEMORY_ID, session_id=session_id, actor_id=actor_id, retrieval_config=retrieval_config), REGION, ) ``` - **Short-term**: Conversation messages stored automatically per session via `AgentCoreMemorySessionManager` - **Long-term**: Strategies extract facts, preferences, summaries and completed episodes asynchronously and make them available as context in future sessions - **Retrieval**: Each namespace has configurable `top_k` and `relevance_score` thresholds for context injection - **Session/Actor IDs**: Derived from AgentCore Runtime request context (`session_id`, `x-amzn-bedrock-agentcore-runtime-user-id` header) - **Actor ID sanitization**: AgentCore Memory requires actor IDs matching `[a-zA-Z0-9][a-zA-Z0-9-_/]*`. Since Cognito emails contain `@` and `.`, `_sanitize_actor_id()` replaces invalid characters with `_` (e.g., `user@example.com` → `user_example_com`). - **Env var**: `MEMORY_SMARTHOMEMEMORY_ID` is auto-set by the `agentcore` CLI on deploy (follows `MEMORY__ID` convention) ### 8.3 Tool Access via AgentCore Gateway (MCP) The Strands Agent discovers and invokes tools through an AgentCore Gateway, which acts as an MCP (Model Context Protocol) server: ``` Strands Agent (MCP Client) | | MCP tool discovery + invocation v AgentCore Gateway (MCP Server) URL: https://{gateway-id}.gateway.bedrock-agentcore.{region}.amazonaws.com/mcp Auth: CUSTOM_JWT (Cognito — user JWT forwarded by agent for per-user policy) Policy Engine: SmartHomeUserPermissions (ENFORCE mode, Cedar permit policies) | | Lambda Target (if permitted by policy) v iot-control Lambda | | Validates command, publishes to IoT Core v AWS IoT Core MQTT ``` The MCP protocol allows the agent to dynamically discover available tools (device control actions) and their schemas from the gateway, then invoke them as needed. #### Tool-argument wrapping for per-user scoping Some backend tools need the caller's identity (KB retrieval, device control, device discovery). AgentCore Gateway's `client_context.custom` does **not** forward the caller's JWT claims to Lambda targets — it only passes Gateway metadata (`bedrockAgentCoreToolName`, target/gateway IDs). So the agent wraps these MCP tools with a thin Strands `@tool` on the client side: 1. The agent decodes the user's `sub` from the runtime-validated idToken once per invocation. 2. For `query_knowledge_base`, `control_device`, and `discover_devices`, it replaces the raw MCP tool with a wrapper of the same short name. The wrapper calls `mcp_client.call_tool_sync(name=)` and injects `user_id=` into the arguments. 3. The LLM sees only the short name and the non-identity arguments. Whatever it tries to pass as `user_id` is overwritten by the wrapper before the Lambda call. Prompt-injection attacks can't escalate to another user's scope. 4. The Lambda refuses any request that arrives without a `user_id` argument, so requests that bypass the wrapper fail fast. Tool suffixes to wrap are matched against both the bare name and the `___` Gateway-prefixed form. #### Strands built-in tools In addition to MCP tools surfaced by the Gateway, the agent registers a small set of Strands `strands_tools` built-ins: - `http_request` — used by the `weather-lookup` skill (Open-Meteo geocode + forecast). - `file_write` — used by the `user-feedback` skill to persist records under `/mnt/workspace/feedback/`. Loaded lazily so the runtime still boots if `strands-agents-tools` is absent. Built-ins default to interactive consent prompts that have no place in a hosted runtime, so `BYPASS_TOOL_CONSENT=true` is set as a runtime environment variable. `agent_core_memory` is **not** auto-registered — it is a provider-style tool (`AgentCoreMemoryToolProvider`) that needs per-session instantiation, and the agent's session manager already persists turns to Memory. ### 8.4 Agent Directory Structure ``` agent/ ├── agent.py # BedrockAgentCoreApp with @app.entrypoint handler ├── vision.py # Claude Haiku 4.5 captioning via Bedrock Converse │ # (used for image-turn bypass; see §8.11) ├── session_storage.py # Per-session filesystem helpers at /mnt/workspace │ # (atomic image writes + index.jsonl catalog) ├── memory/ │ ├── __init__.py │ └── session.py # AgentCoreMemorySessionManager factory (follows agentcore CLI pattern) ├── tools/ │ ├── device_control.py # Fallback tool for local dev (Lambda invocation via boto3) │ ├── browser_use.py # `browse_web` Strands tool — drives AgentCore Browser Tool │ │ # via `browser-use` + DCV (see §9.11) │ └── code_interpreter.py # `execute_python` Strands tool — runs Python in the │ # AgentCore Code Interpreter sandbox (see §9.14) ├── skills/ │ ├── led-control/ # SKILL.md with LED-specific instructions │ ├── rice-cooker-control/ │ ├── fan-control/ │ ├── oven-control/ │ ├── all-devices-on/ # Discovers devices then turns them on sequentially │ ├── weather-lookup/ # Open-Meteo geocoder + forecast via http_request (reference skill) │ ├── user-feedback/ # Persists JSON records under /mnt/workspace/feedback/ via file_write (reference skill) │ ├── browser-use/ # Auto-routes live-web queries to `browse_web` (see §9.11) │ └── code-interpreter/ # Auto-routes data-analysis/charting to `execute_python` (see §9.14) ├── tests/ # pytest unit tests (excluded from CodeZip deploy) ├── pyproject.toml # Dependencies for AgentCore CodeZip packaging └── Dockerfile # Optional, for local container testing ``` The `tests/` directory is intentionally excluded from the CodeZip packaging — `scripts/setup-agentcore.py` filters it out (along with `__pycache__`/`*.pyc`) so the production container never ships test-only imports. **Skills** provide on-demand specialized instructions. Skills are loaded dynamically from DynamoDB per invocation (`load_skills_from_dynamodb`), with the `./skills/` directory serving as a fallback when `SKILLS_TABLE_NAME` is not configured. When the agent activates a skill, it loads the full instructions into context: ```python # Example: agent activates led-control skill # -> Loads instructions about rainbow, breathing, chase modes # -> Knows exact command format: {"action": "setMode", "mode": "rainbow"} # -> Has allowed-tools: control_device ``` > **Tool names in `allowed-tools` must match what the agent actually registers.** > The five device skills originally declared `device_control` while `agent.py` > registers `control_device`; the model dutifully called the declared name, got > `Unknown tool: device_control`, and recovered via a second `skills` call — > a wasted round trip on every device operation that also depressed > `ToolSelectionAccuracy`. Fixed 2026-08-02. ### 8.5 Command Validation The `iot-control` Lambda validates all commands before publishing: ```python DEVICE_COMMANDS = { "led_matrix": { "setMode": {"required": ["mode"], "values": {"mode": [7 modes]}}, "setPower": {"required": ["power"], "types": {"power": bool}}, "setBrightness": {"required": ["brightness"], "ranges": {"brightness": (0, 100)}}, "setColor": {"required": ["color"]}, }, # ... rice_cooker, fan, oven } ``` Invalid commands return 400 with descriptive error messages. The agent receives these errors and can self-correct. ### 8.6 Agent Internal Data Flow The following diagram shows the complete lifecycle of a single user request as it flows through the agent's internal components: ``` HTTP POST /invocations Authorization: Bearer {jwt} {"prompt": "Set LED to rainbow and fan to medium"} | v +-------------------------------+ | BedrockAgentCoreApp | | (bedrock_agentcore) | | | | 1. JWT validation (Cognito) | | 2. Extract request context: | | - session_id | | - actor_id (from header) | +---------------+---------------+ | v +-------------------------------+ | handle_invocation() | | @app.entrypoint | | | | 3. Parse prompt from payload | | 4. Extract session_id, | | actor_id from context | +---------------+---------------+ | v +-------------------------------+ | memory/session.py | | get_memory_session_manager() | | | | 5. Read MEMORY_SMARTHOME- | | MEMORY_ID env var | | 6. Build retrieval config | | per namespace (top_k=3) | | 7. Return SessionManager | +---------------+---------------+ | v +-------------------------------+ | invoke_agent() | | | | 8. Connect MCP client to | | AgentCore Gateway | | 9. Discover tools via MCP | | (list_tools_sync with | | pagination) | | 9a. Load system prompt from | | DynamoDB (per-user + | | __global__ rows, joined | | with "\n\n"; §8.10) — | | fall back to hardcoded | | SYSTEM_PROMPT if both | | rows are empty. | | 10. Create Strands Agent with: | | - BedrockModel (MODEL_ID / | | per-user override) | | - MCP tools | | - Skills plugin | | - Session manager | | - Resolved system_prompt | +---------------+---------------+ | v +-------------------------------+ | Strands Agent | | agent(prompt) | | | | 11. Session manager retrieves | | memory context: | | - /users/{actor}/facts | | - /summaries/{actor}/{sid} | | - /users/{actor}/prefs | | | | 12. LLM receives: | | - System prompt | | - Retrieved memory context | | - Activated skill prompts | | - User prompt | | - Available tool schemas | | | | 13. LLM reasons and emits | | tool_use blocks: | | control_device( | | device_type="led_matrix",| | command={action:"setMode"| | mode:"rainbow"})| | control_device( | | device_type="fan", | | command={action:"setSpeed| | speed:2}) | +------+----------------+-------+ | | +------v------+ +------v------+ | MCP call #1 | | MCP call #2 | | (Gateway) | | (Gateway) | +------+------+ +------+------+ | | +------v------+ +------v------+ | Lambda | | Lambda | | iot-control | | iot-control | +------+------+ +------+------+ | | +------v------+ +------v------+ | IoT Core | | IoT Core | | led_matrix/ | | fan/command | | command | | | +------+------+ +------+------+ | | v v +-------------------------------+ | Device Simulator (Browser) | | MQTT subscribers update UI | +-------------------------------+ | | (back in the agent) v +-------------------------------+ | 14. LLM receives tool results | | and generates final text | | | | 15. Session manager stores | | conversation to memory | | (async extraction runs | | for facts, summaries, | | preferences) | +---------------+---------------+ | v HTTP Response 200 {"response": "Done! I've set the LED matrix to rainbow mode and the fan to medium speed.", "status": "success"} ``` **Key observations:** - Steps 8-9 (MCP tool discovery) happen on every invocation — tools are not cached between requests - Step 11 (memory retrieval) injects relevant context from prior conversations before the LLM sees the prompt - Step 13 (tool calls) may involve multiple sequential or parallel tool invocations depending on the LLM's reasoning - Step 15 (memory storage) is asynchronous — the response is returned before extraction strategies complete - The MCP client connection is scoped to a single invocation (`with mcp_client:` context manager) - Skills are loaded dynamically from DynamoDB per invocation (global + user-specific), with filesystem fallback - The system prompt is also loaded per invocation (global + per-user rows joined with `"\n\n"`), so admin edits take effect on the next request without a runtime redeploy — see [§8.10](#810-agent-system-prompts-text--voice) ### 8.7 Skill Management The agent's skills (specialized instruction sets) are stored in DynamoDB and managed via an admin console. This enables administrators to add, modify, or delete skills without redeploying the agent. **Architecture:** ``` Admin Console (React) | | HTTPS (Bearer JWT) v API Gateway (Cognito Authorizer) | | Lambda Proxy Integration v admin-api Lambda | +--- CRUD operations ---> DynamoDB: smarthome-skills | PK: userId (__global__ or specific user) | SK: skillName | Fields: description, instructions, allowedTools, | license, compatibility, metadata | +--- File management ---> S3: smarthome-skill-files-{accountId} | (presigned URLs) {userId}/{skillName}/scripts/... | {userId}/{skillName}/references/... | {userId}/{skillName}/assets/... | v (per invocation) agent.py: load_skills_from_dynamodb(actor_id) | +-> Query userId = "__global__" (shared skills) +-> Query userId = actor_id (user-specific overrides) +-> Get userId = actor_id, skillName = "__skill_policy__" | -> drop every name in disabledSkills (agent/skill_policy.py) +-> Construct Skill objects (all spec fields) -> AgentSkills plugin v Strands Agent (with dynamic skills) ``` **Per-user skill policy (added 2026-09-18).** The merge above is global UNION per-user, which had no way to take a global skill away from one user — and two global skills are tools, not prose: `browser-use` registers `browse_web` and `code-interpreter` registers `execute_python`, so every user held a live browser and a Python sandbox with no admin lever. The `__skill_policy__` row is the one subtraction: the Skills page, under a user scope, lists the global skills that user inherits with an on/off toggle per skill, saved through `PUT /users/{id}/permissions?action=skills` (same `?action=` dispatch as A2A grants, for the same 20 KB resource-policy reason). The agent applies it after the merge, so switching a tool-bearing skill off removes the tool on the next turn. In the same change `http_request` and `file_write` stopped being registered for everyone unconditionally. They are registered only when an effective skill declares them in `allowedTools` — which the shipped skills already do (`weather-lookup` → `http_request`, `user-feedback` → `file_write`), so nothing a user had disappears, but the tool now follows the skill. `skills is None` (DynamoDB failed, filesystem fallback) keeps the old behaviour: an outage must not become a second, quieter one. **DynamoDB Table Schema (`smarthome-skills`):** All fields from the [Agent Skills specification](https://agentskills.io/specification) are supported: | Attribute | Type | Description | |-----------|------|-------------| | `userId` (PK) | String | `__global__` for shared skills, or Cognito username for per-user | | `skillName` (SK) | String | Skill identifier (e.g., `led-control`). 1-64 chars, lowercase alphanumeric + hyphens | | `description` | String | Skill description, max 1024 chars (shown in skill metadata) | | `instructions` | String | Full markdown instructions (SKILL.md body, loaded on skill activation) | | `allowedTools` | List\ | Tools the skill can use (e.g., `["control_device"]`). Seeded from the SKILL.md `allowed-tools` front-matter, which `seed-skills.py` splits on **commas and whitespace** — splitting on whitespace alone once produced a tool literally named `"discover_devices,"` | | `license` | String | Optional. License name or reference (e.g., `Apache-2.0`) | | `compatibility` | String | Optional, max 500 chars. Environment requirements (e.g., `Requires Python 3.12+`) | | `metadata` | Map\ | Optional. Arbitrary key-value pairs for additional metadata | | `createdAt` | String | ISO 8601 timestamp | | `updatedAt` | String | ISO 8601 timestamp | **Skill File Storage (S3):** Skill directory files (scripts, references, assets) are stored in S3: ``` S3: smarthome-skill-files-{accountId} {userId}/{skillName}/scripts/extract.py {userId}/{skillName}/references/REFERENCE.md {userId}/{skillName}/assets/template.json ``` - Files are managed via presigned URLs (browser uploads/downloads directly to S3) - Admin API generates time-limited presigned URLs for PUT (upload) and GET (download) - Only three directories are allowed: `scripts`, `references`, `assets` - When a skill is deleted, all its S3 files are cascade-deleted **Skill Override Rule:** When a user-specific skill has the same name as a global skill, the user-specific version takes precedence. Global skills are loaded first, then user-specific skills overwrite by name. **Fallback:** If the `SKILLS_TABLE_NAME` environment variable is not set or the DynamoDB query fails, the agent falls back to loading skills from the local `./skills/` directory (the 5 built-in skills). **Admin API Endpoints:** | Method | Path | Description | |--------|------|-------------| | `GET` | `/skills?userId=` | List skills for a user scope. With `&promptBundle=1` returns the Agent Prompt bundle (text + voice) instead — see [§8.10](#810-agent-system-prompts-text--voice). | | `POST` | `/skills` | Create a new skill (all spec fields) | | `GET` | `/skills/{userId}/{skillName}` | Get a single skill (or prompt record when `skillName` matches `__prompt_text__` / `__prompt_voice__`) | | `PUT` | `/skills/{userId}/{skillName}` | Update a skill, or upsert a prompt override when `skillName` starts with `__prompt_` | | `DELETE` | `/skills/{userId}/{skillName}` | Delete a skill (cascade-deletes S3 files), or remove a prompt override when `skillName` starts with `__prompt_` | | `GET` | `/skills/users` | List distinct user scopes | | `GET` | `/skills/{userId}/{skillName}/files` | List files in a skill directory | | `POST` | `/skills/{userId}/{skillName}/files/upload-url` | Get presigned S3 upload URL | | `POST` | `/skills/{userId}/{skillName}/files/download-url` | Get presigned S3 download URL | | `DELETE` | `/skills/{userId}/{skillName}/files?path=` | Delete a skill file | | `GET` | `/settings/{userId}` | Get user settings (`modelId` + `visionModelId`) | | `PUT` | `/settings/{userId}` | Update user settings (either field, independently patchable) | | `GET` | `/sessions` | List all user runtime sessions | | `GET` | `/users` | List all Cognito users with groups | | `GET` | `/tools` | List all gateway tools (reads tool schemas from targets) | | `GET` | `/users/{userId}/permissions` | Get user's allowed tools | | `PUT` | `/users/{userId}/permissions` | Update allowed tools + sync Cedar policies | | `GET` | `/memories` | List all memory actors | | `GET` | `/memories/{actorId}` | Get long-term memory records (facts + preferences) for an actor | | `GET` | `/dashboard?range=24h\|7d\|30d` | Ops-dashboard fast half — health metrics, evaluation scores, release/version state (see [§9.15](#915-agent-operations-dashboard)) | | `GET` | `/dashboard?range=…&part=spans&dim=user\|tenant\|agent` | Ops-dashboard slow half — TTFT percentiles and token trend/attribution from a Logs Insights query over the runtime span groups (~5-20s; see the note in §9.4) | **Authorization:** All admin API endpoints require a valid Cognito JWT. The Lambda additionally checks that the caller belongs to the `admin` Cognito group (returns 403 if not). > `/dashboard` is wired with a plain `apigw.Integration` plus a **single** > wildcard `lambda:InvokeFunction` permission rather than CDK's per-method > auto-permissions, because the admin Lambda's auto-generated resource policy is > already at the API Gateway 20 KB cap — the same reason `/optimization/*` and > the `?action=` dispatches exist. ### 8.7.1 Prompt Caching Every orchestrator turn re-sends an identical ~10.5k-token prefix: the system prompt (~1.6k), the A2A routing table (~1.7k), eleven governed skills (~4.3k) and ~20 tool schemas. `create_agent` passes `cache_config=CacheConfig(strategy="auto")` so Bedrock reads it from cache instead of reprocessing it. Measured on the live runtime: | | `InputTokenCount` | `CacheReadInputTokenCount` | `CacheWriteInputTokenCount` | |---|---|---|---| | before | 29,644 | 0 | 0 | | after | **9** | 20,988 | 10,512 | **This is a cost optimisation, not a latency one, and the AWS documentation reads otherwise.** The docs advertise "up to 85% latency reduction"; measured side-by-side at ~15.6k prefix tokens over two controlled runs (12 and 8 alternating calls), latency moved 2% — 2.11s uncached against 2.06s cached, and 2.70s against 2.68s. At this prefix size the prefill was never the bottleneck. The 98% token cut is real and every turn of every user pays that prefix, so it is worth having; calling it a latency win would be a claim the numbers do not support. `strategy="auto"` rather than a hardcoded cache point because the model is per-user configurable (§8.8): auto asks Strands to check model support and degrade to uncached with a warning, where `cache_prompt="default"` would fail every turn for a user on an unsupported model. Cache hits need an **exact** prefix match, which is why the static prompt/skills/tools come first and per-request content (the governed override, the user's memory, the JSON-output rules) is appended last. The A2A specialists deliberately do **not** cache — see [Delegation latency, measured](#delegation-latency-measured) in §9.13 for why. The same cost-not-latency lesson is stated as a principle in [agent-design-principles §2.5](agent-design-principles.md#25-cache-what-repeats-and-know-what-caching-buys). **These numbers are Anthropic-specific.** They were measured with explicit Anthropic cache points via `CacheConfig(strategy="auto")` on `bedrock-runtime`, and the default model (`us.anthropic.claude-sonnet-4-6`) is on that path, so they still describe it. They would NOT describe a Bedrock Mantle model: caching there is automatic prefix caching with no cache point to place. Mantle is not integrated (§8.8), but if it is re-added, this table needs re-measuring rather than inheriting. `create_agent` also logs `prompt caching: strands= model= strategy=<...>` once per agent build. That line exists because the failure is silent in both directions: caching that quietly stops working looks identical to caching that works, and diagnosing it from outside cost an hour — the spans carry no cache attributes, so a high `InputTokenCount` with an absent `CacheReadInputTokenCount` was the only signal, and it is indistinguishable from "the code was never deployed". The live line reads `strands=1.51.0 strategy=anthropic`; note 1.51.0, where `requirements.txt` pins only `>=1.25.0`, so the container's resolved version was itself unknown until it said so. ### 8.8 Per-User Model Selection Administrators can assign different LLM models to each user (or set a global default) via the Admin Console for **both** the text agent and the vision agent. Settings are stored in DynamoDB using a reserved key `skillName = "__settings__"`. **Storage format:** ``` PK: userId (e.g., "zihangh@amazon.com" or "__global__") SK: "__settings__" modelId: "us.anthropic.claude-sonnet-4-6" # text agent model (a cross-region profile id) modelEndpoint: "runtime" # which Bedrock endpoint serves modelId visionModelId: "us.anthropic.claude-haiku-4-5-..." # vision (multimodal) model ``` `PUT /settings/{userId}` accepts either field independently — absent fields retain their existing values so callers can patch one without clobbering the other. `modelEndpoint` is written by the console alongside `modelId`, resolved from the live catalog at the moment the admin picks the model (see below). **Resolution order (text):** user-specific `modelId` > `__global__` `modelId` > `MODEL_ID` env var (default: `us.anthropic.claude-sonnet-4-6`). **Resolution order (vision):** user-specific `visionModelId` > `__global__` `visionModelId` > `VISION_MODEL_ID` env var (default: `us.anthropic.claude-haiku-4-5-20251001-v1:0`). The agent reads this via `load_user_settings(actor_id)` and threads it into `vision.caption_images(..., model_id=...)`. **Available text models** are no longer a hardcoded list. `GET /settings/{userId}?action=catalog` on the admin API builds it live from two `bedrock-runtime` calls: | Call | Why both | |---|---| | `bedrock:ListFoundationModels` | The models, with `inputModalities` (vision) and `modelLifecycle` (deprecation) | | `bedrock:ListInferenceProfiles` | The id you can actually **invoke**. A newer Claude model rejects its bare id with "Invocation of model ID ... with on-demand throughput isn't supported"; the cross-region profile id (`us.`, `global.`) is the invocable one | **One row per model.** Where a base model has a system-defined profile the profile id wins and the bare id is dropped: offering both would show two entries differing only in routing, and the bare one is the one that fails at invoke time. Where there is no profile, the bare id is offered if it supports `ON_DEMAND`. Live on this deployment that yields **88 models**. This replaced 33 hand-maintained entries in `AdminConsole.tsx` that had to be edited every time Bedrock shipped a model. The response carries a `catalogError` string: when a listing fails the console renders a warning rather than an empty picker, because a failed lookup and a genuinely empty account are otherwise the same empty dropdown — the same reasoning as the A2A catalog in §9.9. `modelEndpoint` is stored next to `modelId` and always resolves to `runtime` today. It exists because the agent cannot call the admin API, so a second endpoint would need the answer stored at selection time rather than re-derived on every cold start. **Bedrock Mantle is deliberately not integrated.** The `bedrock-mantle` endpoint is the only way to reach the GPT-5.x family, and a working two-path version of this was built and verified against the live endpoint before being removed on request. The findings are kept in `agent/model_provider.py` so re-adding it does not repeat the discovery: - The listing is `https://bedrock-mantle.{region}.api.aws/v1/models`. The `/openai/v1` path shown in the Gemma 4 blog and the GPT-5.6 Luna model card **404s** for listing. - Mantle serves two OpenAI-compatible APIs on different paths and **no model accepts both**: `openai.gpt-5.6-luna` and `google.gemma-4-31b` are Responses-only (`/openai/v1`), `minimax.minimax-m2.5` is Chat-Completions-only (`/v1`). The wrong one returns `400 The model '...' does not support the '...' API`. - Nothing reports which format a model accepts — `/v1/models` and `/v1/models/{id}` return status and data-retention only. It has to be probed and cached. - Auth is a short-term bearer token (`aws-bedrock-token-generator`), not SigV4, and needs `bedrock-mantle:CreateInference|Get*|List*` plus `CallWithBearerToken`. **Vision models** are the same catalog filtered to `inputModalities` containing `IMAGE` — 44 of the 88 on this deployment. ### 8.9 Per-Login Session ID and Session Tracking Each page load gets a **fresh runtime session ID** generated once in the chatbot (`chatbot/src/components/ChatInterface.tsx` — `loginSessionIdRef`). The same id is reused by warmup, text chat, voice, and browser-use for the whole login so they all land on the same AgentCore Runtime microVM (the post-login warmup heats that microVM, and subsequent chat turns skip the ~5s cold start). Refreshing or re-logging-in rotates the id. **Session ID format:** `user-session-{cognito-sub}-{Date.now()}` (set by the chatbot via the `X-Amzn-Bedrock-AgentCore-Runtime-Session-Id` header). The timestamp suffix scopes each AgentCore **Online Evaluation** batch to a single login's traces — without it, historical traces from previous days accumulate under the same `session.id` and a single broken Strands span poisons every subsequent evaluation with `AgentSpanMappingException: Failed to parse user_query`. Cross-session continuity is instead preserved by AgentCore Memory's long-term summaries (keyed on `actor_id = cognito sub`, not `session_id`), so STM resets per login but LTM carries context across logins. **User identification:** The AgentCore Runtime strips the `X-Amzn-Bedrock-AgentCore-Runtime-User-Id` header before forwarding to the agent process. To work around this, the chatbot passes the user's email in the POST body as `userId`, and the agent reads it from `payload.get("userId")`. **Session tracking:** On each invocation, the agent records `{userId, sessionId, kind, lastActiveAt}` to a dedicated DynamoDB table (`smarthome-runtime-sessions`, PK `userId`, SK `sessionKey = "{kind}#{sessionId}"`). Each per-login session gets its own row — no overwrites — so the Admin Console's Sessions tab can render the full per-user history. The `kind` attribute is `text` for text-runtime rows (written by `agent._record_session`) and `voice` for the voice runtime (written by `voice_session._record_voice_session`); the UI's `Kind` column maps to the correct runtime ARN when stopping (`POST /sessions/{id}/stop?kind=text|voice`). This table is independent of the `smarthome-skills` table — runtime sessions and skill storage are unrelated concerns. **Per-session 7-day token usage:** The Sessions tab also shows each session's total token consumption over the last 7 days. Each `chat` span (emitted by `strands.telemetry.tracer`) carries both `attributes.session.id` and `attributes.gen_ai.usage.total_tokens`. The admin Lambda runs a CloudWatch Logs Insights query on `GET /sessions` that sums `total_tokens` grouped by `session.id` for the last 7 days and joins the result onto the DynamoDB session rows as `totalTokens7d`. The query is the backing dataset for the CloudWatch "GenAI Observability → Bedrock AgentCore → All sessions" dashboard, so the numbers match what an admin sees in that console view. > ⚠️ **Where the spans live changed, silently.** AgentCore Runtime used to export > Strands/ADOT spans to the account-wide `aws/spans` log group. Since **2026-08-05** > (orchestrator) and **2026-08-09** (the eight A2A specialists) each runtime writes > them to its own group instead — a `spans` stream inside > `/aws/bedrock-agentcore/runtimes/{runtimeId}-DEFAULT`. > > The cutover was clean and completely invisible: `aws/spans` stops at 02:12 and > the runtime-local stream starts at 02:28 the same day. Nothing errored, no > permission was denied, and `StartQuery` kept succeeding — it simply matched zero > records, so the Overview page's TTFT and token cards and this tab's token totals > read "no data" **for six days**, exactly as they would on an idle system. > > `dashboard._spans_log_groups()` now queries **both** sources and merges them: a > 30d dashboard range still reaches back past the cutover, so dropping the legacy > group would erase five weeks from a trend line the page exists to draw. > > Two things had to be measured rather than assumed: > - **`StartQuery` rejects the WHOLE request** with `ResourceNotFoundException` if > any one named group is missing, so one torn-down specialist runtime still > listed in `DASHBOARD_EXTRA_RUNTIME_ARNS` would take down every spans card. The > group list is filtered through `DescribeLogGroups` first > (`_existing_log_groups`), and a describe failure leaves the list intact — > losing the ability to check is not evidence that nothing exists. > - **The `POST /invocations` span is under the `opentelemetry.instrumentation.starlette` > scope**, not the Strands one. Filtering on the Strands scope alone silently > drops it, and it is the only measurement of how much of a request's wall clock > was spent inside this container. Permissions: the admin Lambda gets `logs:StartQuery`/`logs:StopQuery` on **both** `log-group:aws/spans:*` and `log-group:/aws/bedrock-agentcore/runtimes/*`, plus `logs:DescribeLogGroups` and `logs:GetQueryResults` (both required at `*` — neither supports resource-level scoping). The runtime groups are wildcarded rather than enumerated because the specialists' runtime ids belong to `a2a-agent-registry/deploy.py`, not to the CDK stack: naming them would mean a stack deploy per specialist redeploy, with a silently failing read until it happened. **Stop session:** The Admin Console calls the AgentCore `StopRuntimeSession` API from the browser using the admin's AWS credentials (obtained by exchanging the Cognito idToken for Identity Pool temporary credentials — the same SigV4 flow the chatbot uses for `/invocations` and `/ws`). `scripts/setup-agentcore.py` also invokes this API as a post-deploy step to invalidate any warm sessions so users pick up fresh code immediately instead of waiting for the idle timeout. ### 8.10 Agent System Prompts (Text & Voice) Administrators can override the system prompts used by both the **text agent** (Strands on HTTP `/invocations`) and the **voice agent** (BidiAgent on WebSocket `/ws`) from the Admin Console's **Agent Prompt** tab, without redeploying the runtime image. **Two independent records, additive at runtime.** The global prompt and per-user prompt are stored as two separate DynamoDB rows and the agent concatenates them at invocation time: ``` effective_system_prompt = global_body + "\n\n" + user_body ``` Any empty part is omitted from the join. When both rows are empty, the agent falls back to the hardcoded `SYSTEM_PROMPT` / `VOICE_SYSTEM_PROMPT` constant shipped in the container image. This keeps shared guardrails in one editable place (Global) while letting per-user personalization be a short addendum rather than a full duplicate. **Storage format** (reuses the existing `smarthome-skills` table): | userId | skillName | Fields | |--------|-----------|--------| | `__global__` or `{cognito-email}` | `__prompt_text__` | `promptBody`, `updatedAt`, `updatedBy` | | `__global__` or `{cognito-email}` | `__prompt_voice__` | `promptBody`, `updatedAt`, `updatedBy` | `updatedBy` captures the admin's email from the Cognito JWT on each save, surfaced as the "Last edited by …" line in the UI. **Agent runtime resolution** (`agent/agent.py:load_system_prompt`): 1. Read `(userId=actor_id, __prompt_{type}__)` — returns `""` if row missing. 2. Read `(userId=__global__, __prompt_{type}__)` — returns `""` if row missing. 3. Concatenate non-empty parts with `"\n\n"`. If both empty, return `None` → caller uses hardcoded constant. The text agent calls this inside `invoke_agent()` on every request; the voice agent calls it inside `handle_voice_session()` after decoding the JWT `actor_id`. Both paths silently fall back to the hardcoded constant on DynamoDB errors (`logger.warning` + default) — a broken prompt-load must never break the invocation. **Why the voice prompt has to stay editable separately.** The voice prompt isn't a stripped-down copy of the text one — it uses the MCP-gateway-prefixed tool names (e.g., `SmartHomeDeviceDiscovery___discover_devices`) that Nova Sonic needs to emit in `toolUse` events, plus voice-specific tone rules ("one short spoken sentence, no Markdown"). Forcing admins to edit them together would make it easy to break voice by assuming the text prompt's shorthand tool names still work. **Admin UI.** At Global scope the tab shows two side-by-side editors (Text / Voice). At per-user scope each editor splits into three rows: a read-only Global base view, an editable per-user addendum, and a collapsible "Effective Prompt Preview" that concatenates them the same way the agent does. Badges indicate the state (`Built-in Default`, `Custom Global`, `No User Override`, `User Override`), and the reset button is labelled `Revert to Default` at global scope or `Remove Override` at per-user scope, so the semantic difference between "wipe everyone's custom prompt" and "drop this user's addendum" is never ambiguous. **Transport (no new API Gateway methods).** The admin API carries prompts on the **existing** `/skills` routes to stay under the admin Lambda's 20 KB resource-policy cap. The CDK passes `allowTestInvoke: false` to the admin `LambdaIntegration` which suppresses AWS's per-method `/test-invoke-stage/*` permission (halving the policy — 20 KB → ~10 KB — so new routes and `browse_web`-related dispatch fit comfortably). Even with that headroom, the prompt API piggybacks on `/skills` and the browser-session-poll API piggybacks on `/sessions` via `?action=` dispatch because adding methods is still expensive and a fresh cold deploy would otherwise be within a few hundred bytes of the cap. Requests are dispatched to prompt handlers when `skillName` starts with the reserved `__prompt_` prefix: | Method | Path | Behavior when skillName matches `__prompt_*__` | |--------|------|------------------------------------------------| | `GET` | `/skills?userId={scope}&promptBundle=1` | Returns `{text: PromptRecord, voice: PromptRecord, userId}` where each record = `{body, updatedAt, updatedBy, isOverride, globalBody, builtinDefault}`. `globalBody` is always populated (at user scope it's the read-only context; at global scope it equals `body`). | | `PUT` | `/skills/{userId}/__prompt_text__` | Upsert the text prompt override; rejects empty or >16 KB bodies. | | `PUT` | `/skills/{userId}/__prompt_voice__` | Same, for voice. | | `DELETE` | `/skills/{userId}/__prompt_{type}__` | Remove the override; agent falls back to the next resolution level on the next invocation. | **Every agent's prompt, not just these two.** `PROMPT_AGENT_TYPES = ("text", "voice")` is now `BUILTIN_AGENT_TYPES`, and any **AgentCard name** is also a valid `agentType` — so `__prompt_light-effect-agent__` governs that specialist exactly as `__prompt_text__` governs the orchestrator. The valid set is derived from the Registry's approved AGENT records rather than hardcoded, and re-read on a miss before rejecting: a stale cache is precisely how the Cedar action map once silently authorised nothing, and an agent deployed after the container warmed up would otherwise be permanently unaddressable with a 400 that points at the caller. Sub-agent prompts are edited from the agent detail page (§9.18) and resolved by the runtime per request (§9.13). **Defaults mirror.** The Lambda ships `agent_prompt_defaults.py` (text + voice, hand-maintained) and `a2a_prompt_defaults.py` (the eight specialists, **generated** by `scripts/gen-prompt-defaults.py` from each `system_prompt.md`). The duplication is intentional: the admin Lambda is packaged from its own directory and cannot read the agent tree, and making the editor render "what the agent will use when no override exists" without a round-trip is worth the copy. A stale mirror is *worse* than a missing one — "Revert to Default" would install a prompt no agent has ever run, and the editor would misreport what an override changes. Nothing else catches it: the mirror is a valid Python string whatever it says, and the agent never reads it. `test_prompt_defaults_mirror.py` compares all nine against their sources byte for byte. The text/voice mirror was previously kept in sync by a comment asking future editors to remember; that comment is now enforced. **Cross-reference to AgentCore Optimization.** Each editor card renders an info Alert "Optimization Suggestions (AgentCore Optimization)" with a link "Open in Optimization tab →" that deep-links to `#/optimization`. The optimization workflow (recommendations, configuration bundles, A/B tests) is documented in §8.12 — applying a recommendation writes back into the same `__prompt_text__` / `__prompt_voice__` rows this section describes, so the two tabs share storage and an applied recommendation immediately becomes the effective prompt for subsequent invocations. **Image-awareness clause (text prompt).** The hardcoded text `SYSTEM_PROMPT` contains an `IMAGES IN THIS CONVERSATION:` section that tells the text model image descriptions are injected into the conversation history as prior assistant messages (written by the vision bypass path in §8.11). Without this clause the text model would respond to follow-up questions about past uploads by denying image access or fabricating contents. Admins editing the global prompt must keep this section when overriding, or follow-up turns referring to earlier images will regress. --- ### 8.11 Image Input (Vision Bypass Path) Image turns use a dedicated path that bypasses the text model entirely: the runtime calls a vision model (Claude Haiku 4.5) to caption the uploaded images and returns the description straight to the chatbot. This keeps latency low (one model call instead of two) and keeps the text model's tool-calling / KB / skills loop focused on text. ``` POST /invocations {"prompt": "...", "userId": "...", "images": [{"mediaType": "image/png", "data": ""}]} │ ▼ handle_invocation (agent/agent.py) │ ├── payload.images present? ── no ──► invoke_agent() → text model │ (normal text path) │ └── yes: vision bypass │ ├─ validate: isinstance(list) && len ≤ 3 → 400 on violation │ ├─ 1. PERSIST raw bytes to /mnt/workspace//uploads/images/ │ via session_storage.save_image (atomic tempfile + os.replace, │ dedup by sha256, append index.jsonl entry). Non-fatal on error. │ ├─ 2. vision.caption_images(images, prompt) │ → bedrock-runtime.converse on VISION_MODEL_ID │ (default us.anthropic.claude-haiku-4-5-20251001-v1:0; │ env-overridable, e.g. us.amazon.nova-lite-v1:0). │ → Per-image validation (MIME allowlist, ≤20 MB after decode) │ is defense-in-depth mirroring the client cap. │ → Partial-success: rejected images emit a "Note: …" warning; │ ≥1 valid image → the call proceeds. │ → One retry on Throttling/ServiceUnavailable; on final │ failure returns a placeholder + service-unavailable note. │ ├─ 3. PERSIST exchange to AgentCore Memory (short-term event): │ (user_prompt, USER) + (haiku_description, ASSISTANT), each │ wrapped in the Strands SessionMessage envelope so the text │ agent's session manager can deserialize them on the next text turn. │ Metadata carries a fingerprint per image (mime/bytes/sha16) │ and image_N_path pointing to the session-storage file. │ └─ 4. Return {"response": caption_text + warnings, "status": "success"} ``` **Why bypass the text model.** Measured with a local latency harness (not shipped in the repo; 30 rounds × 2 models × 3 image counts), the bypass path lands at p50 ≈ 3.3 s (Haiku, 1 image, cold) / 2.0 s (Nova Lite, 1 image, cold). Routing the same turn through the text model after pre-captioning would roughly double that (two LLM hops plus an extra token-pricing cost on the text model for the long description). The measurement was taken when the text model was Kimi K2.5. **Model swap.** `VISION_MODEL_ID` is captured at module-load (so each container version is pinned to one vision model) and resolved via Bedrock `Converse` for whichever multimodal model the operator prefers. Tested paths: Claude Haiku 4.5 (default — stronger OCR / fine-detail) and Amazon Nova Lite (~1.7× faster, lower cost, lighter detail). Swapping is an `update_agent_runtime` env-var patch — no source-code change and no `agentcore deploy`. The update rolls a new container version; subsequent sessions start on the new model. **Memory envelope.** Because the Strands `AgentCoreMemoryConverter` writes and reads events as JSON-serialized `SessionMessage` dicts, the vision bypass must mirror that format exactly. Plain-text events cause `JSONDecodeError` in `list_messages()` and silently drop the entire short-term history for the session, leaving the text model with no memory of uploaded images. `agent/agent.py:_persist_vision_turn` serializes each message as `{"message": {"role": "...", "content": [{"text": "..."}]}, "message_id": N, "redact_message": null, "created_at": "...", "updated_at": "..."}` so later text turns round-trip cleanly. **Per-session filesystem.** The runtime is configured with `filesystemConfigurations=[{sessionStorage: {mountPath: "/mnt/workspace"}}]` (set by `scripts/setup-agentcore.py`). Every `runtimeSessionId` gets an isolated writable volume at `/mnt/workspace//` that survives across invocations within the same session. The layout: ``` /mnt/workspace// └── uploads/ └── images/ ├── __. # raw bytes, atomic writes └── index.jsonl # append-only catalog ``` `index.jsonl` rows: `{ts, sha256, mime, bytes, path, user_prompt, caption_event_id?}`. The directory is reserved for the image subsystem today; the top-level `uploads/` namespace is intentional so future modalities (audio, documents) can coexist without churn. The volume is automatically torn down when the session is closed by AgentCore — no TTL sweep needed. **Client-side limits (defense mirrored server-side):** | Limit | Value | Rationale | |-------|-------|-----------| | Max images per turn | 3 | Fits AgentCore metadata cap (1 fingerprint + 1 path per image + `image_count` = 7 of 15 KV slots) | | Max bytes per image | 20 MB raw | Fits well inside the 100 MB runtime payload cap even at the 3-image × base64 bloat worst case | | MIME allowlist | `png \| jpeg \| webp \| gif` | Matches Bedrock Converse `image/format` support | **Error surfaces:** | Failure | Behavior | |---------|----------| | Client: >3 files / bad MIME / too large | Inline error under thumbnail strip; not sent | | Server: `images` not a list or >3 items | 400 `{"error": "Invalid images payload (max 3)."}` | | Server: per-image bad MIME or decoded >20 MB | Soft-reject in `caption_images`, warning lists indices, remaining images still captioned | | Vision model throttled / unavailable after one retry | Placeholder caption + `Note: vision service was unavailable…`, user still gets a reply | | Memory write fails after successful caption | Caption returned to user anyway; next turn simply won't see this one in memory (logged at WARN) | --- ### 8.12 AgentCore Optimization (Recommendations + Target-Based A/B Routing) > **Per-agent targets (added with the Agents page).** The optimisation target > used to be a two-option segmented control (`text` / `tool_desc`), and > `_runtime_arn_for` returned the **text** runtime for every non-voice value — so > naming a sub-agent would have analysed the orchestrator's traces and produced a > recommendation for the wrong prompt. The target is now a selector populated > from the fleet read model (§9.18), so every deployed specialist appears without > a frontend change, and the ARN is resolved from the fleet too. An unknown id > resolves to `""` rather than the orchestrator's ARN, because analysing the wrong > agent while appearing to succeed is worse than failing. Administrators drive a data-driven prompt-improvement loop from the Admin Console's **Assess → Optimization** tab without redeploying the runtime image. The tab orchestrates AgentCore Optimization (public preview) **Recommendations** and **target-based A/B Tests** for the text agent's system prompt, plus the legacy **Configuration Bundle** path for tool descriptions. The original implementation was config-bundle-based (variants = different bundle versions on the same runtime, agent reads bundle via W3C baggage). It is replaced for prompts by the **target-based** pattern from the official AgentCore docs: variants reference different runtime-endpoint qualifiers via dedicated gateway targets. Target-based covers all change types (prompt, model, tool descriptions, code) and matches the chatbot's normal HTTP `/invocations` flow. See `docs/superpowers/specs/2026-05-17-agentcore-optimization-target-based-design.md` for the full design doc. ``` Admin Console (React) AWS ┌─────────────────────────────────────────────┐ ┌──────────────────────────────────┐ │ Assess → Optimization │ │ bedrock-agentcore (data plane) │ │ ┌────────────────────────────────────────┐ │ │ start/get/list/delete │ │ │ A/B Routing toggle (Global) [● ─ ──] │ │ │ _recommendation │ │ │ Scope · AgentType: Text │ ToolDescs │ │ │ create/get/list/update │ │ │ Recommendations │ │ │ _ab_test │ │ │ Tool Description Bundles (when ToolD) │ │ ── ▶│ │ │ │ A/B Tests (when Text · disabled OFF) │ │ │ bedrock-agentcore-control │ │ │ ToolDesc-A/B-not-automated Alert │ │ │ create/get/list/update/delete │ │ └────────────────────────────────────────┘ │ │ _agent_runtime_endpoint │ └─────────────────────────────────────────────┘ │ _gateway, _gateway_target, │ │ │ _online_evaluation_config, │ ▼ /optimization/* (one wildcard perm) │ _configuration_bundle* │ ┌─────────────────────────────────────────────┐ └──────────────┬───────────────────┘ │ admin-api Lambda optimization.py │ ▲ │ ab-toggle (NEW): GET/PUT __opt_routing_… │ ◀────── DDB ───── │ │ recs · target-based ab-tests · bundles │ smarthome-skills │ │ apply: text→__prompt_text__ only │ │ │ tool_desc→UpdateGatewayTarget+bundle│ │ └─────────────────────────────────────────────┘ │ │ Chatbot (text path) │ ┌─────────────────────────────────────────────┐ │ │ ChatInterface → SigV4 POST │ │ │ {gw-opt}/smarthome-control/invocations │ │ │ (always, regardless of toggle state — the │ │ │ gateway routes by gatewayFilter + │ │ │ active A/B test, not by URL) │ │ │ Voice path stays direct → wss://runtime/ws │ │ └──────────────────┬──────────────────────────┘ │ ▼ │ ┌─────────────────────────────────────────────┐ │ │ Optimization Gateway (NEW, dedicated) │ │ │ smarthome-optimization-gateway-{id} │ │ │ authorizerType=AWS_IAM │ │ │ Targets: │ │ │ smarthome-control → endpoint "control" │ │ │ smarthome-treatment → endpoint "treatment"│ │ │ No active test → 100% to smarthome-control │ │ │ Active test → sticky session-id split │ │ │ per CreateABTest weights │ │ └──────────────────┬──────────────────────────┘ │ ▼ │ ┌─────────────────────────────────────────────┐ │ │ smarthome runtime (existing) │ │ │ 2 named endpoints on the same runtime: │ │ │ control / treatment (initially same ver) │ │ │ Each endpoint logs to its own log group: │ │ │ /aws/.../{rt-id}-control │ │ │ /aws/.../{rt-id}-treatment │ │ │ Both read __prompt_text__ at request time │ ◀──────────────────┘ └─────────────────────────────────────────────┘ Tools Gateway (existing — UNCHANGED) smarthome-smarthomegateway-{id}, authorizerType=CUSTOM_JWT Targets: SmartHomeDeviceControl, …Discovery, …KnowledgeBase Both runtime endpoints call this for MCP tool calls. Tool-description recommendations + bundles still target this gateway via UpdateGatewayTarget. ``` **Why a separate optimization gateway.** The tools gateway uses `CUSTOM_JWT` for per-user Cedar policy. The optimization gateway needs `AWS_IAM` because the chatbot signs `bedrock-agentcore` SigV4 (matches the runtime's existing IAM auth flow), and target-based A/B routing applies to that surface. Running both auth modes on one gateway isn't supported. Lifecycle isolation is the secondary win — tool integrations don't churn when an experiment starts or stops. **Why "always go through the gateway."** Single chatbot code path. The toggle's effect lives entirely server-side: when OFF the gateway routes 100 % to `smarthome-control`; when ON the gateway splits per the active A/B test. Steady-state cost: ~50 ms gateway hop in exchange for zero chatbot routing complexity. **Voice path is unchanged.** AgentCore Gateway proxies HTTP `/invocations` only, not WebSocket `/ws`. Voice mode keeps signing the existing `wss://bedrock-agentcore.{region}.amazonaws.com/runtimes/{voice-arn}/ws` URL directly to the voice runtime. Voice prompts stay editable from the Agent Prompt tab; they just can't participate in A/B traffic splitting. **Five moving pieces.** 1. **Admin UI tab** — Cloudscape `Container`s top-to-bottom: A/B Routing toggle (global) → Scope+AgentType filter (`text` | `tool_desc`) → Recommendations → Tool Description Bundles (only when `tool_desc`) → A/B Tests (only when `text`, button disabled when toggle OFF). Tool-desc view also renders a non-dismissible Alert "A/B testing is not automated yet" explaining why MCP tool descriptions can't be A/B-tested through the existing AgentCore primitives. 2. **Admin Lambda module** `cdk/lambda/admin-api/optimization.py` — handler groups `start/list/get/delete/apply_recommendation`; `list/create/get/delete_bundle` (gated to `tool_desc`); `start/list/get/stop_ab_test` (target-based); `get/put_ab_toggle` (NEW). Module-level preview-availability probe (`PREVIEW_UNAVAILABLE`) returns 501 `AgentCoreOptimizationUnavailable` from every handler so the UI shows a banner instead of a confusing 5xx. Top-level dispatch wraps `/optimization/*` in try/except → `optimization._resp(500, …)` so unexpected exceptions still carry CORS headers. 3. **DDB rows** in the existing `smarthome-skills` table under reserved sort keys, partitioned by scope (`__global__` or user email): - `__opt_rec_{id}__` — recommendation index row. - `__opt_bundle_{arn}__` — configuration-bundle mirror (now tool_desc only). - `__opt_abtest_{id}__` — A/B test index row, including target-based fields (`routingMode`, `controlEndpoint`, `treatmentEndpoint`, `controlTarget`, `treatmentTarget`). - `__opt_routing_enabled__` — global A/B routing toggle. Deliberately NOT under the `__opt_abtest_` prefix to avoid collision with the `abtest` opt-row kind in `_query_opt_rows("__global__", "abtest")`. 4. **Dedicated optimization gateway + runtime endpoints**, provisioned at deploy time by `setup-agentcore.py:_ensure_optimization_infra()`: 2 named runtime endpoints (`control`, `treatment`) on the existing smarthome runtime, both initially pinned to the latest version; the gateway with `authorizerType=AWS_IAM`; 2 AgentCore-runtime gateway targets (`smarthome-control`, `smarthome-treatment`); 2 per-endpoint online-eval configs (`smarthome_control_online_eval`, `smarthome_treatment_online_eval`) each scoring **6 builtin evaluators** — `Builtin.GoalSuccessRate` + `Builtin.Helpfulness` (defaults from the helper) plus `Builtin.Correctness`, `Builtin.InstructionFollowing`, `Builtin.ToolSelectionAccuracy`, `Builtin.Conciseness` (added operationally for richer A/B signal). 3 supporting IAM roles (`smarthome-abtest-execution-role`, `smarthome-optimization-gateway-role`, `smarthome-optimization-online-eval-role`). All steps idempotent. 5. **No more runtime baggage hook for prompts.** `agent/bundle_config.py` is retained but text-prompt apply no longer creates a bundle; both endpoints simply read `__prompt_text__` at request time. The bundle path stays alive only for `tool_desc` (rollback snapshot — `UpdateGatewayTarget` is the actual mechanism, the bundle is a versioned record of the pre-apply tool-desc payload). **Dedicated `/optimization/*` REST subtree, one wildcard permission.** The CDK uses plain `apigw.Integration` (no auto-emitted `AWS::Lambda::Permission`) plus a single `addPermission` with `sourceArn=…/*/*/optimization/*`. Net policy growth: ~800 bytes total, flat regardless of how many methods we add — relevant because the admin Lambda's resource-based policy is already near the 20 KB cap. **API surface** (all under Cognito JWT auth, admin group required): | Method + Path | Behavior | |---|---| | `GET` / `PUT /optimization/ab-toggle` | Read or set the global routing flag. PUT body `{enabled: bool}`. **Side effect on PUT OFF**: looks up any active `__opt_abtest_*__` row (executionStatus in `NOT_STARTED`/`RUNNING`/`PAUSED`), calls `UpdateABTest(executionStatus="STOPPED")`, finalizes the DDB mirror, and returns `stoppedTestId` in the response. **Side effect on PUT ON**: none — gateway already exists; no test starts until admin clicks Start A/B Test. | | `POST /optimization/recommendations` | Start a recommendation. Body: `{scope, agentType, evaluatorArn, startTime, endTime, name?, ruleFilter?, logGroupArn?, serviceName?}`. When `logGroupArn` is omitted the Lambda builds the **per-runtime** log group ARN (`/aws/bedrock-agentcore/runtimes/{rt-id}-DEFAULT` for text, `…{voice-rt-id}-DEFAULT` for voice). When `serviceName` is omitted it's derived from the runtime ARN as `{runtime_short}.DEFAULT`, matching the value AgentCore stamps on every runtime span. Returns `{recommendationId, recommendationArn, status}` (202). | | `GET /optimization/recommendations?scope=…` | List from DDB, cached status. | | `GET /optimization/recommendations/{recId}` | Live `get_recommendation`; surfaces `recommendedSystemPrompt` / `tools[]` when COMPLETED, refreshes DDB row. | | `POST /optimization/recommendations/{recId}/apply` | **text/voice**: writes the optimized prompt to `__prompt_{type}__` only — no bundle. Both runtime endpoints (control + treatment) read this row at request time so the change takes effect on the very next invocation. To A/B test "new prompt vs old prompt", the admin must first deploy a new runtime version and repoint the `treatment` endpoint to it via the agentcore CLI; the prompt itself can't differ between endpoints. **tool_desc**: unchanged — calls `UpdateGatewayTarget` on the tools gateway and snapshots a configuration bundle for rollback. | | `DELETE /optimization/recommendations/{recId}` | `delete_recommendation` + DDB cleanup. | | `GET /optimization/bundles` / `POST` / `GET /{bundleArn}` / `DELETE /{bundleArn}` | List filtered server-side to `agentType=tool_desc`; `POST` rejects other agentTypes with 400. The legacy text/voice `__opt_bundle_*__` rows from the pre-redesign deploy stay in DDB but are invisible to the UI. | | `POST /optimization/ab-tests` | Refuses when toggle is OFF (400 `ToggleDisabled`); rejects `voice`/`tool_desc` agentTypes (400 `UnsupportedAgentType`); rejects per-user scope (400). Body: `{agentType: "text", controlEndpoint, treatmentEndpoint, variantWeights, durationDays}`. Single-active rule per `agentType`. The handler builds a target-based `CreateABTest`: `gatewayArn=OPTIMIZATION_GATEWAY_ARN`, `gatewayFilter.targetPaths=["/smarthome-control/*"]`, `evaluationConfig.perVariantOnlineEvaluationConfig=[{name:"C",arn:CONTROL_…},{name:"T1",arn:TREATMENT_…}]`, variants reference `target.name=smarthome-control`/`smarthome-treatment`. `enableOnCreate=True` so the test starts running immediately. | | `GET /optimization/ab-tests` / `GET /{testId}` / `POST /{testId}/stop` | List / live results (per-variant mean, p-value, winner, CloudWatch dashboard deep-link, `routingMode`, `controlEndpoint`, `treatmentEndpoint`) / stop. Stop → `update_ab_test(executionStatus="STOPPED")` + follow-up `get_ab_test` to compute the final winner. | **Live-API constraints discovered during implementation.** The official boto3 model enforces several regexes and field shapes that the spec didn't anticipate: - **Variant names must match `(C|T1)`** (≤2 chars). Friendlier names like `"control"` / `"treatment"` 400 with `ValidationException`. The handler hardcodes `C` and `T1`. - **A/B test names regex `[a-zA-Z][a-zA-Z0-9_]{0,47}`** — no hyphens. Auto-name uses `abtest_{agent_type}_{YYYYMMDD_HHMMSS}`. - **Online-eval config names same regex** — `smarthome_control_online_eval` / `smarthome_treatment_online_eval` (underscores, not hyphens). - **`components` for `CreateConfigurationBundle` is a dict, not a list.** Shape: `{component_arn: {"configuration": {...}}}`. The earlier list-of-dicts produced `ParamValidationError` at boto3. - **`CreateGatewayTarget` for AgentCore-runtime targets needs `credentialProviderConfigurations=[{credentialProviderType: "GATEWAY_IAM_ROLE"}]`.** Same shape as MCP Lambda targets — without it the call fails with "Credential provider configurations is not defined". - **Gateway `CreateGateway` returns immediately but the gateway is in `CREATING` state for ~10–30s.** Calling `CreateGatewayTarget` during that window 400s with "Cannot perform operation … when gateway is in CREATING status". The setup helper polls `get_gateway` until `status` leaves `CREATING`. - **Chatbot SigV4 against `https://{gw}.gateway.bedrock-agentcore…/{target}/invocations` requires `bedrock-agentcore:InvokeGateway`** on the Cognito auth role — `InvokeAgentRuntime` alone returns 403. The setup script now grants both actions on the optimization gateway ARN. - **Recommendation API needs the per-runtime log group**, not `aws/spans`. The IAM template in the official docs grants `logs:StartQuery,GetQueryResults,FilterLogEvents,GetLogEvents` on `arn:aws:logs:*:*:log-group:/aws/bedrock-agentcore/runtimes/*` (note the leading slash). The admin Lambda's `logs:StartQuery` resource list was extended to cover `/aws/bedrock-agentcore/runtimes/*` in addition to the legacy `aws/spans` ARN. - **Toggle-row SK collision.** Originally the toggle was stored at `__opt_abtest_enabled__`, which `_query_opt_rows("__global__","abtest")` matched as `begins_with(__opt_abtest_)` and surfaced as a phantom test. Renamed to `__opt_routing_enabled__`. Teardown deletes both old and new SKs for migration safety. **`apply_recommendation` behavior — by agent type.** - **text / voice** — writes only the `__prompt_{type}__` DDB row. No bundle. Both `control` and `treatment` runtime endpoints pick it up on the next invocation (they read the same DDB row, no caching). Returns `{applied: true, agentType, scope}`. - **tool_desc** — unchanged. Calls `UpdateGatewayTarget` on the **tools gateway** for each recommended tool, then snapshots the full `tools` list as a single bundle component (component ARN is the literal string `"tool_desc"`). The bundle exists for rollback only — there is no AgentCore-side per-target A/B routing path that would consume it. Failures on individual gateway-target updates are warn-and-continue. Returns `{appliedBundleArn, appliedBundleVersionId}`. **Why tool-description A/B testing is not automated.** The Optimization tab renders an info Alert in the tool-desc view explaining the gap: AgentCore A/B routing applies at the **agent runtime layer** (target-based, switching runtime endpoints) or via **bundles read by the agent runtime** (config-bundle, BeforeModelCallEvent hook). Neither path reaches MCP tool-description metadata served by a tools gateway. Automating it would require either two complete tools-gateway surfaces (two parallel sets of Lambda integrations) — which collapses back into runtime-level A/B — or per-request dual-MCP-client logic in the agent code, which is custom application logic, not an AgentCore feature. The current workflow for tool-desc is "apply → observe → roll back via the Tool Description Bundles snapshot if needed." **Storage layout** (existing `smarthome-skills` DynamoDB table, reused the same way `__prompt_*__` reuses it): | userId (PK) | skillName (SK) | Fields | |---|---|---| | `__global__` | `__opt_routing_enabled__` | `value: bool`, `updatedAt`, `updatedBy` | | `__global__` or `{email}` | `__opt_rec_{id}__` | `recommendationArn`, `agentType`, `status` (cached), `evaluatorArn`, `logGroupArn`, `startTime`, `endTime`, `createdAt`, `createdBy`, `appliedAt?`, `appliedBundleVersionId?`, `__opt_expires?` (TTL) | | `__global__` | `__opt_abtest_{id}__` | `abTestArn`, `agentType`, `routingMode: "target-based"`, `controlEndpoint`, `treatmentEndpoint`, `controlTarget`, `treatmentTarget`, `variantWeights`, `controlOnlineEvalArn`, `treatmentOnlineEvalArn`, `durationDays`, `autoStopAt`, `status` + `executionStatus` (cached), `createdAt`, `createdBy`, `stoppedAt?`, `winner?`, `__opt_expires?` (auto-stop + 30 days) | | `__global__` | `__opt_bundle_{arn}__` | tool_desc only — `bundleArn`, `bundleName`, `latestVersionId`, `agentType: "tool_desc"`, `sourceRecommendationId?`, timestamps | **Status caching rule.** List views show DDB-cached status. Detail views call `get_recommendation` / `get_ab_test` live and write the fresh status back. Same pattern as the Sessions tab. Winners are computed locally from `results.evaluatorMetrics[].variantResults[].isSignificant + mean > controlStats.mean` because the AgentCore API does not return a `winner` field directly. **Chatbot config injection** (`scripts/setup-agentcore.py` writes these to `chatbot/config.js` post-deploy): ```js window.__CONFIG__ = { ..., agentRuntimeArn: "arn:…runtime/smarthome_smarthome-…", // kept for voice WSS voiceAgentRuntimeArn: "arn:…runtime/smarthomevoice_…", // voice WSS direct optimizationGatewayUrl: // NEW — text path "https://smarthome-optimization-gateway-{id}.gateway.bedrock-agentcore.{region}.amazonaws.com", optimizationDefaultTarget: "smarthome-control", // NEW }; ``` `chatbot/src/voice/sigv4.ts` exports `signedGatewayInvocationsFetch({gatewayUrl, targetName, …})` for the gateway path; `signedInvocationsFetch` is retained for the warmup ping (which warms the runtime microVM directly). `ChatInterface.tsx` picks the gateway helper when `optimizationGatewayUrl` is set, otherwise falls back to the runtime helper (transitional safety for old `config.js`). #### 8.12.1 Operating an A/B test — running playbook The setup deploys the optimization infrastructure but stops short of producing **different** behavior between the two endpoints. To run an A/B test that actually compares two configurations, four practical findings from live runs apply: **1. The two endpoints start identical.** `setup-agentcore.py` creates both `control` and `treatment` runtime endpoints pinned to the same `agentRuntimeVersion` (whatever was current at deploy time). An A/B test against this state runs but every evaluator returns p≈1.0 because the variants are byte-identical. To produce a real signal: edit `agent.py` (e.g. change `SYSTEM_PROMPT`, `MODEL_ID`, or any code path), run `agentcore deploy` from `.agentcore-project/smarthome/` to get a new runtime version, then `UpdateAgentRuntimeEndpoint(endpointName="treatment", agentRuntimeVersion=)` to repoint **only** treatment. Control stays on the old version. Revert the local `agent.py` afterward so future deploys don't accidentally promote the treatment prompt to baseline. **2. `agentcore deploy` resets several runtime fields**, then auto-bumps the version when you call `UpdateAgentRuntime` to restore them. Specifically, `environmentVariables`, `requestHeaderConfiguration` (the `X-Amzn-Bedrock-AgentCore-Runtime-Custom-AuthToken` allowlist), `protocolConfiguration`, `authorizerConfiguration`, and `filesystemConfigurations` are all stripped by `agentcore deploy` and must be reapplied via `update_agent_runtime`. Each `update_agent_runtime` call creates a **new** version, so the deploy → restore sequence produces v_N (deploy) and then v_{N+1} (restore). Pin treatment to v_{N+1}, since v_N is missing the env vars/header allowlist and would 401 against the gateway. The setup script handles this for the initial deploy via `_ensure_optimization_infra`; ad-hoc treatment redeploys must replay the same restore logic. **3. The DDB `__prompt_text__` row overrides the runtime's hardcoded `SYSTEM_PROMPT`.** Per §8.10, the agent reads `(__global__, __prompt_text__)` and `({user}, __prompt_text__)` at every invocation and concatenates them with the hardcoded constant. If a previous Apply Recommendation step wrote a `__global__` `__prompt_text__` row, **both** runtime versions read that row instead of their own baked-in prompt — meaning the A/B variants again behave identically regardless of which version each endpoint is pinned to. To isolate the runtime-version difference for an A/B test, delete the `__global__` `__prompt_text__` row (save its contents first if you want to restore later) before driving traffic. Alternatively, edit it so it carries instructions you want to test. Since the introduction of `ab-bundles` mode (§8.13), this row-deletion workaround is only needed for `ab-targets`-mode tests; bundles-mode tests leave it untouched. **4. Online-eval results need ≥ sessionTimeoutMinutes + ~10–15 min after the last session.** AgentCore aggregates results only after each session crosses the configured `sessionTimeoutMinutes` idle window AND after the eval pipeline catches up. The default `sessionTimeoutMinutes=5` adds a 5-minute floor; lowering it to **1 minute** shrinks the wait for small-scale tests. A common pitfall: stop the test too early (e.g. ~5 min after last session) and the `results` field is permanently `None` because the aggregator never runs to completion against a STOPPED test. The minimum safe stop time is `last_session + sessionTimeoutMinutes + 10 minutes`. `setup-agentcore.py` provisions both per-variant online-eval configs at `sessionTimeoutMinutes=5` by default; reducing them is a deliberate operator action via `update_online_evaluation_config`. **5. Image-only turns are unevaluable.** Vision-bypass (§8.11) calls `bedrock-runtime.converse` directly without going through Strands, so no `strands.telemetry.tracer` `chat` span is emitted for that turn. The online-eval pipeline filters by "supported scope names" and silently records a `ValidationException: No spans with supported scope names found for traces: []` for any session whose only chat-span is the vision call. Drive **text-only** prompts when running an A/B test if you want every session counted. **6. The runbook in practice** (the steps below are the whole of it — there is no separate runbook file): ``` 1. Lower sessionTimeoutMinutes on both eval configs (5 → 1): update_online_evaluation_config(rule.sessionConfig.sessionTimeoutMinutes=1) 2. Edit agent.py SYSTEM_PROMPT (or MODEL_ID, or behavior) → treatment variant. 3. Sync agent/ → .agentcore-project/smarthome/app/smarthome/, then `agentcore deploy` from .agentcore-project/smarthome/. Note new version v_N. 4. Snapshot runtime config; call update_agent_runtime to replay env vars + requestHeaderConfiguration + filesystemConfigurations. New version v_{N+1}. 5. update_agent_runtime_endpoint(name=treatment, agentRuntimeVersion=v_{N+1}). Verify: control.liveVersion < treatment.liveVersion. 6. Revert agent.py so the local source matches control behavior. 7. (Targets mode only.) If you want to isolate runtime-version differences, ensure no __global__ __prompt_text__ row would shadow the difference. **Prefer using `ab-bundles` mode (§8.13) when the only thing you want to compare is global prompt text** — bundles bypass DDB at the model-call level via a `BeforeModelCallEvent` hook, so production prompt rows stay intact during the test. 8. Spot-check: invoke {gw}/smarthome-control vs {gw}/smarthome-treatment with the same prompt + different session-ids. Outputs must visibly differ. 9. PUT /optimization/ab-toggle {enabled:true} (if not already on). 10. POST /optimization/ab-tests with desired weights + durationDays. 11. Drive N text-only sessions, ≥70 s apart, fresh session-id each. 12. Wait (last-session-time + 16 min) before reading results or stopping. 13. POST /optimization/ab-tests/{id}/stop. Read winner from response. ``` **Sample run** (from this branch): n=30, control v212 (verbose baseline) vs treatment v217 (brevity-tuned prompt; identical otherwise): | Evaluator | Control μ (n=19) | Treatment μ (n=11) | Δ% | p-value | sig | |---|---|---|---|---|---| | Conciseness | 0.053 | **0.864** | **+1540.9%** | **0.0000** | ✓ | | Correctness | 0.947 | 1.000 | +5.6% | 0.591 | ✗ | | GoalSuccessRate | 0.947 | 1.000 | +5.6% | 0.757 | ✗ | | Helpfulness | 0.823 | 0.830 | +0.8% | 0.982 | ✗ | | InstructionFollowing | 0.947 | 0.818 | -13.6% | 0.533 | ✗ | | ToolSelectionAccuracy | 1.000 | 0.969 | -3.1% | 0.604 | ✗ | `POST /ab-tests/{id}/stop` returned `{executionStatus: "STOPPED", winner: "T1"}` — admin Lambda's `_compute_winner` correctly picked T1 because at least one evaluator (Conciseness) had `isSignificant=True` and `treatment.mean > control.mean`. With identical prompts (the misconfigured baseline) all six evaluators had `p > 0.5` and winner stayed `null` regardless of sample size. ### 8.13 Per-Tenant Entry Environment Each tenant has an `entryEnvironment` selecting which AgentCore surface the chatbot connects to. Stored in DDB as `(__global__, __tenant_env_{email}__) → {mode, updatedAt, updatedBy}`. Missing row resolves to `default`. | Mode | Chatbot URL | Server-side prompt source | Per-user prompt | A/B dimension | |---|---|---|---|---| | `default` | runtime SigV4 (primary) | DDB additive (§8.10) | ✓ | None | | `ab-bundles` | runtime SigV4 (bundles runtime) | `BeforeModelCallEvent` hook reads bundle from baggage, bypasses DDB | ✗ (masked — UI confirms) | Global prompt text | | `ab-targets` | optimization gateway → control/treatment endpoints | DDB additive on each endpoint | ✓ | Runtime version (model + code + prompt) | **`smarthome_bundles` belongs to `ab-bundles`, not to `ab-targets`.** Worth stating explicitly because the name invites the opposite reading, and it has been misread: the runtime named after "bundles" is the one **target-based** A/B never touches. Target-based A/B routes through the optimization gateway to two *endpoints* on the **primary** runtime — verified live, both targets carry `arn:...runtime/smarthome_smarthome-{id}` with qualifiers `control` and `treatment`. Nothing in that path names the bundles runtime. Renaming it was considered and declined. `agentRuntimeName` is immutable, so a rename is create-plus-delete: a new `runtimeId`, every stored `bundlesRuntimeArn` re-pointed (`config.js`, the admin Lambda env), and — because spans moved to per-runtime log groups (§9.x) — historical traces and token attribution split across the old and new names at the cutover. The accurate name would be `smarthome_promptbundle`; the cost of getting there is not worth paying for a name, so the documentation carries the correction instead. **Two runtimes, one image.** Primary runtime (`smarthome_smarthome-{id}`) and bundles runtime (`smarthome_bundles-{id}`) share an ECR image. The underscore in `smarthome_bundles` is mandatory — AgentCore's `agentRuntimeName` regex is `[a-zA-Z][a-zA-Z0-9_]{0,47}`, so hyphens 400 on `CreateAgentRuntime`. They differ only in the `ENABLE_BUNDLE_HOOK` env var: when `=1`, `agent.py:load_system_prompt` returns None (no DDB) and `create_agent` registers a Strands `BeforeModelCallEvent` hook that reads the W3C `baggage` header and overrides `system_prompt` on each model call. When unset, DDB additive resolution per §8.10. **Why two runtimes.** The hook registers globally on the Strands `Agent`; there is no per-request "skip" path. One runtime would have to either always-hook (breaking per-user prompts) or never-hook (breaking the bundles-A/B path). Two runtimes keeps each surface single-purpose. **Admin Console.** "Entry Environment" section above the A/B Routing toggle on the Optimization tab. The Add/edit modal's tenant field is a **Cloudscape Select with autosuggest filtering** populated from `listCognitoUsers()` (the `GET /users` endpoint shared with the Identity tab) — admins pick from the existing user pool rather than free-typing emails, which prevents typos that would create dangling override rows for non-existent users. In Add mode the dropdown hides users that already have an override (use Edit instead); in Edit mode the field is locked to the row being edited. Switching to `ab-bundles` while a per-user `__prompt_text__` row exists produces a 409 `PerUserPromptWillBeMasked`; the UI shows a confirm dialog and re-submits with `acknowledgeMaskedOverride: true`. **Chatbot.** `ChatInterface.tsx` calls `getTenantMode(userEmail)` (60-sec TTL cache) before each send and selects the URL accordingly. Cache miss or fetch failure falls back to `default`. Mode changes propagate within 60 seconds; no forced reconnect, no WebSocket push. **API surface** (piggybacks on `/skills` to stay under the admin Lambda's 20 KB resource-policy cap — same trick §8.10 uses for `__prompt_*__`): | Method + Path | Behavior | |---|---| | `GET /skills?tenantEnv=1` | List all overrides. | | `GET /skills?tenantEnv=1&userId={email}` | Single tenant's mode (default if no row). | | `PUT /skills/{email}/__tenant_env__` | Upsert; 409 if `ab-bundles` would mask an existing per-user prompt and `acknowledgeMaskedOverride` is not set. | | `DELETE /skills/{email}/__tenant_env__` | Remove override → tenant falls back to `default`. | **Voice path is unchanged.** The mode applies only to text `/invocations`. Voice keeps its WSS direct-to-runtime path. --- ## 9. Infrastructure Design ### 9.1 Two-Stack Architecture The system deploys as two CloudFormation stacks: **Stack 1: `SmartHomeAssistantStack`** (managed by CDK) — standard AWS resources: ``` Cognito User Pool +-> User Pool Client + Domain +-> Identity Pool + Auth/Unauth IAM Roles +-> Admin Group (for skill management access) IoT Endpoint (Custom Resource) +-> iot-control Lambda (env: IOT_ENDPOINT) — validates & publishes MQTT commands +-> iot-discovery Lambda — returns the device catalog (shared/device-catalog.json) +-> iot-query Lambda — current device state + sensor history (query_device_state, query_sensor_history) +-> iot-history-ingest Lambda — IoT rule target that stores reported state/history +-> nav-deeplink Lambda — navigate_to_page deep links +-> Device Simulator config.js DynamoDB Table (smarthome-skills) +-> Skill storage (global + per-user, full Agent Skills spec fields) +-> Admin Lambda read/write access S3 Bucket (smarthome-skill-files) +-> Skill directory files: scripts/, references/, assets/ +-> CORS enabled for presigned URL uploads from browser +-> Admin Lambda read/write access Admin API (API Gateway + Lambda) +-> Cognito Authorizer (admin group check) +-> CRUD endpoints for skill management (all spec fields) +-> File management endpoints (presigned URL upload/download, list, delete) S3 Buckets + CloudFront (x4) +-> Device Simulator, Chatbot, Admin Console, Skill ERP +-> BucketDeployment (static assets) +-> Custom Resource (config.js injection) ``` **Stack 2: `AgentCore-smarthome-default`** (managed by `agentcore` CLI) — AgentCore resources: ``` AgentCore Gateway (MCP Server) +-> Auth: CUSTOM_JWT (Cognito — same User Pool as runtime) +-> Policy Engine: SmartHomeUserPermissions (ENFORCE mode) +-> Lambda Target: SmartHomeDeviceControl (iot-control Lambda + tool schema) +-> Lambda Target: SmartHomeDeviceDiscovery (iot-discovery Lambda + tool schema) +-> Lambda Target: SmartHomeKnowledgeBase (kb-query Lambda + tool schema) +-> Lambda Target: SmartHomeDeviceQuery (iot-query Lambda; added by setup-agentcore.py via API) +-> Lambda Target: SmartHomeNavigation (nav-deeplink Lambda; added by setup-agentcore.py via API) AgentCore Runtime +-> CodeZip (Python 3.14, agent code from S3) +-> BedrockAgentCoreApp (Strands Agent) +-> Auth: AWS_IAM (SigV4; agentcore.json carries no JWT authorizer) +-> requestHeaderAllowlist: ["X-Amzn-Bedrock-AgentCore-Runtime-Custom-AuthToken"] | (carries the user idToken to the agent; `Authorization` cannot be allowlisted under AWS_IAM) +-> Env: AGENTCORE_GATEWAY_{NAME}_URL, MEMORY_SMARTHOMEMEMORY_ID, MODEL_ID, SKILLS_TABLE_NAME AgentCore Memory (managed by agentcore CLI: `agentcore add memory`) +-> SEMANTIC strategy (fact extraction) +-> SUMMARIZATION strategy (session summaries) +-> USER_PREFERENCE strategy (preference learning) +-> EPISODIC strategy (ordered account of what happened) +-> Env var auto-set: MEMORY_SMARTHOMEMEMORY_ID ``` Created outside that stack by `setup-agentcore.py` / `a2a-agent-registry` scripts via the control-plane API: ``` smarthome-websearch-gw (us-east-1, CUSTOM_JWT) -> Target: SmartHomeWebSearch (web-search connector) smarthome-a2a-gw (scripts/setup-a2a-gateway.py) -> one target per A2A specialist runtime (§9.13) Optimization gateway + bundles runtime (A/B entry environments, §8.12 / §8.13) ``` **Why two stacks?** AgentCore resources cannot be created via CDK for two reasons: 1. The CDK JS SDK does not yet include `@aws-sdk/client-bedrockagentcorecontrol`, so `AwsCustomResource` fails 2. A Python Lambda custom resource also fails because the Lambda runtime's bundled boto3 is too old for the `bedrock-agentcore-control` API The `agentcore` CLI solves both by using its own CDK stack with the native `AWS::BedrockAgentCore::Runtime` CloudFormation type and packaging agent code as CodeZip (no Docker required). **Key `agentcore` CLI quirks discovered during deployment:** - `agentcore create --defaults` must be used (not `--no-agent`) to create a deploy target - `aws-targets.json` must be seeded with `[{"name": "default", "region": "...", "account": "..."}]` for non-interactive deploy - `agentcore deploy -y --verbose` enables non-interactive mode (the default TUI requires a TTY) - `--outbound-auth` is not applicable for `lambda-function-arn` targets (Lambda targets don't need credential provider config) - Agent code must include `pyproject.toml` for the CLI to package it as CodeZip - Gateway `authorizerType` cannot be changed after creation — must delete and recreate the CloudFormation stack - Gateway auth should be `CUSTOM_JWT` for per-user tool control. The runtime is `AWS_IAM` (SigV4) because CUSTOM_JWT on the Runtime's `/ws` endpoint is not reliable today; the chatbot forwards the idToken in the custom header `X-Amzn-Bedrock-AgentCore-Runtime-Custom-AuthToken` (allowlisted via `requestHeaderAllowlist` in `UpdateAgentRuntime`), and the agent re-wraps it as `Bearer` on the gateway MCP client. `Authorization` **cannot** be allowlisted under AWS_IAM — `UpdateAgentRuntime` rejects it. - The `agentcore` CLI sets gateway URL env vars as `AGENTCORE_GATEWAY_{GATEWAYNAME}_URL` (not `AGENTCORE_GATEWAY_URL`); agent code must auto-detect the pattern - `agentcore deploy` drops custom `environmentVariables` set in `agentcore.json` — must patch them post-deploy via `update_agent_runtime` boto3 API (requires passing `agentRuntimeArtifact`, `roleArn`, `networkConfiguration`, and `authorizerConfiguration` alongside). CLI-managed env vars like `MEMORY__ID` and `AGENTCORE_GATEWAY__URL` are preserved. - **`requestHeaderConfiguration` round-trip pitfall.** `get_agent_runtime` returns the allowlist as the top-level field `requestHeaderAllowlist`, but `update_agent_runtime` expects it nested under `requestHeaderConfiguration={"requestHeaderAllowlist": [...]}`. A naïve round-trip that re-passes the top-level value **silently drops the allowlist**, which strips the custom auth header at the edge proxy → the agent sees `context.request_headers = None` → MCP gateway returns 401. Any helper that calls `UpdateAgentRuntime` (latency-probe nonce bumpers, welcome-message toggles, session redeploy helpers) must read from either location and always wrap back into the nested form. `scripts/restore-text-runtime-config.py` is the in-repo example of the correct pattern. - `agentcore add memory --name --strategies SEMANTIC,SUMMARIZATION,USER_PREFERENCE,EPISODIC` adds memory as a project resource deployed via the same CFN stack; the CLI auto-sets `MEMORY__ID` env var on the runtime - ⚠️ **`.agentcore-project/smarthome/app/smarthome/` is a COPY of `agent/`, not a link, and `agentcore deploy` packages whatever is sitting in it.** Only `setup-agentcore.py` refreshes that copy (`shutil.copytree`, as one step of a full provision), so the normal way to ship an agent change — edit `agent/`, run `agentcore deploy` — deploys the code from the *last full provision* and reports success. Hit on 2026-08-11 with two features at once: both were committed, both "deployed", the runtime went READY on a new version, and neither was in the container — the copy was five hours old and did not even contain the new files. Every symptom pointed elsewhere, and both plausible diagnoses were wrong: the specialists ignoring a prompt (which had genuinely happened an hour earlier for a different reason), and prompt caching not working through AgentCore. So an orchestrator change is three commands, not one: `sync-agent-code.py`, `agentcore deploy`, `restore-text-runtime-config.py`. The exact sequence and the manual diff check are in the admin manual [§2 (§2.5)](admin_manual_管理员使用手册.md#2-agent-代码快速部署); `sync-agent-code.py --check` exits 1 on a stale copy, so a deploy wrapper or CI can fail rather than ship one. - ⚠️ **A warm container keeps running the old code.** Independent of the above: the runtime updated at 06:36 and a 06:38 turn still executed the previous version, showing zero cache activity on correctly deployed code. Force a **new** `runtimeSessionId` to get a cold container before concluding anything from a post-deploy test. `setup-agentcore.py` stops DynamoDB-tracked sessions for this reason; an ad-hoc `agentcore deploy` does not. - **`restore-text-runtime-config.py` is not optional.** `agentcore deploy` strips the A2A/table env vars, `protocolConfiguration`, the header allowlist and the `/mnt/workspace` mount. Without the A2A vars the orchestrator registers **no** `a2a_*` tools and answers every specialist question itself — silently. Measured on real deploys during this work: 12 vars restored each time. The m2m vars (`A2A_M2M_SECRET_ARN`, `A2A_COGNITO_*`) went away with the user-idToken model (§9.13); the script's `env_wanted` is the current list, so check by name, not by count. ### 9.2 Deployment Architecture `deploy.sh` is a thin wrapper that runs 7 split scripts under `scripts/0[1-7]-*.sh` in order. Each split script is independently runnable and prints the AWS resources it creates at the top — handy for debugging or re-running a single step after a partial failure. The diagram below is what each step builds and why it is ordered that way; the operator procedures (minimal redeploy, the three commands a manual `agentcore deploy` needs) are in the admin manual [§2](admin_manual_管理员使用手册.md#2-agent-代码快速部署), and the end-to-end demo walkthrough is in [agentcore-deploy-runbook.md](agentcore-deploy-runbook.md). ``` deploy.sh (one-click wrapper) | +---> [1/7] scripts/01-install-deps.sh | npm install in cdk/; upgrade boto3 in the venv (Registry API needs | >= 1.42.93); bundle latest boto3 into Lambda dirs (admin-api, | user-init, kb-query, skill-erp-api) for AgentCore control-plane APIs | +---> [2/7] scripts/02-build-frontends.sh | Build device-simulator, chatbot, admin-console, skill-erp React apps | +---> [3/7] scripts/03-cdk-bootstrap.sh | cdk bootstrap (idempotent; CDKToolkit stack, asset bucket, ECR) | +---> [4/7] scripts/04-cdk-deploy.sh | cdk deploy --all | CloudFormation: SmartHomeAssistantStack | -> Cognito (User Pool, Identity Pool, admin group, admin user) | -> IoT Things + endpoint | -> Lambda (iot-control, iot-discovery, iot-query, iot-history-ingest, | nav-deeplink, admin-api, kb-query, user-init, pre-token, | skill-erp-api, scenario-runner) | -> DynamoDB smarthome-skills | -> S3 (smarthome-skill-files, smarthome-kb-docs) | (S3 Vector bucket + index created later in step 6 by setup-agentcore.py) | -> Bedrock Knowledge Base + S3 data source (Cohere multilingual) | -> API Gateway with Cognito authorizer | -> S3 + CloudFront for each of the 4 frontends (incl. Skill ERP) | -> Outputs written to cdk-outputs.json | +---> [5/7] scripts/05-fix-cognito.sh | aws cognito-idp update-user-pool | -> AllowAdminCreateUserOnly=false (self-service sign-up) | -> auto-verified-attributes=email | (CDK selfSignUpEnabled doesn't always propagate reliably) | +---> [6/7] scripts/06-deploy-agentcore.sh -> scripts/setup-agentcore.py | | | +---> agentcore create --name smarthome --defaults | +---> Replace default agent code with agent/ | | Patch agentcore.json (entrypoint, env vars; no JWT authorizer -> AWS_IAM) | | Seed aws-targets.json (required for CLI deploy) | +---> agentcore add memory --name SmartHomeMemory | | --strategies SEMANTIC,SUMMARIZATION,USER_PREFERENCE,EPISODIC | +---> agentcore add gateway (CUSTOM_JWT auth, Cognito) | +---> agentcore add gateway-target SmartHomeDeviceControl | | (iot-control Lambda + control_device tool schema) | +---> agentcore add gateway-target SmartHomeDeviceDiscovery | | (iot-discovery Lambda + discover_devices tool schema) | +---> agentcore add gateway-target SmartHomeKnowledgeBase | | (kb-query Lambda + query_knowledge_base tool schema) | +---> (after deploy, via control-plane API) | | gateway-target SmartHomeDeviceQuery (iot-query Lambda: | | query_device_state, query_sensor_history) | | gateway-target SmartHomeNavigation (nav-deeplink Lambda: | | navigate_to_page) | | smarthome-websearch-gw (us-east-1) + target SmartHomeWebSearch | +---> agentcore add evaluator + online-eval | +---> agentcore deploy -y --verbose | | CloudFormation: AgentCore-smarthome-default | | -> Runtime, Gateway, Memory, Policy Engine, IAM Role | +---> Fetch resource IDs (from CFN stack outputs) | +---> Initialize KB (S3 Vector bucket + index, Bedrock KB, S3 data source, | | default __shared__/ + admin@/ folders) | +---> Patch runtime env vars (MODEL_ID, AWS_REGION, SKILLS_TABLE_NAME) | +---> Patch runtime: AWS_IAM auth + requestHeaderAllowlist: | | ["X-Amzn-Bedrock-AgentCore-Runtime-Custom-AuthToken"] | | (carries the user idToken to the agent for gateway auth) | +---> Patch admin Lambda env vars (AGENT_RUNTIME_ARN, GATEWAY_ID, | | SKILL_FILES_BUCKET, COGNITO_USER_POOL_ID) | +---> Patch user-init Lambda env vars (GATEWAY_ID, KB_DOCS_BUCKET) | +---> Grant runtime role DynamoDB + Bedrock Retrieve access | +---> Update config.js in S3 for device-sim / chatbot / admin console | | (admin config.js also injects chatbotUrl + deviceSimulatorUrl | | so the Tool Access tab can deep-link per-user demo flows) | +---> Invalidate CloudFront cache | then scripts/set-platform-version.py --apply | -> text, voice and bundles runtimes V1 -> platform V2 (§9.2.1) | +---> [7/7] scripts/07-seed-skills.sh -> scripts/seed-skills.py +---> Read SKILL.md files from agent/skills/ +---> Write to DynamoDB as __global__ skills (idempotent) ``` ### 9.2.1 Runtime Platform V2 (Snapshot Start) All 11 runtimes (text, voice, bundles and the eight A2A specialists) run on AgentCore Runtime `platformVersion: V2`, moved from V1 on 2026-10-01. V2 starts the container once per runtime version, takes a snapshot on the first healthy `/ping`, and restores every new session's microVM from that snapshot instead of re-running the import. **What it bought, measured.** Both runs used `scripts/measure-baseline.py --repeats 2`: the same 10 prompts, a fresh session each, and the rows whose span join failed excluded. The `platform` column is the time spent before our container is entered. | | V1 | V2 | |---|---|---| | `platform`, fresh session | 7.3-10.8s (fast group mean 8.3s) | 2.4-3.7s (fast group mean 2.8s) | | fast-group wall, first turn | 18.7s | 15.6s | Cold start is shorter, not gone: a first turn still waits about 3s for the session and then for the LLM. The first session after a new version goes READY ran about 11s longer inside the container (`harness` 15.8s against the usual 4-5s), and later sessions did not. **Setting it.** `scripts/set-platform-version.py` (`--apply`, `--only `, `--to V1` to roll back). Neither CloudFormation, CDK nor the agentcore CLI can set the field, so a runtime the CLI creates is always V1. Three facts follow from that: - An `UpdateAgentRuntime` that omits `platformVersion` keeps the current one. So `agentcore deploy`, `restore-text-runtime-config.py` and the setup script's env patches all leave V2 alone. - A fresh provision does not, so `06-deploy-agentcore.sh` runs the script after `setup-agentcore.py` (text, voice, bundles). `a2a-agent-registry/deploy.py` declares `PLATFORM_VERSION` in its post-deploy patch for the specialists. The bundles runtime is created on the primary's platform so the two A/B arms differ only in `ENABLE_BUNDLE_HOOK`. - `platformVersion` is recorded per runtime **version**. The `control` (v212) and `treatment` (v217) endpoints of the text runtime are pinned to versions that predate V2 and still run V1. A V2 copy of old code cannot be made without moving `DEFAULT`, because every update becomes the new `DEFAULT`. Both arms are V1, so an experiment across them is still like for like. **Limits that now bite.** A V2 update takes 4-5 minutes to go READY (measured 244-274s, against seconds on V1), and a second update on the same runtime before then fails with `ConflictException`. The env cap is 1.5 KB for CodeZip (V1: 4 KB). The largest runtime here is bundles at about 1.0 KB of JSON, so this is the cap a new env var will hit first. The container must report healthy within 120s of start; the import takes 2-4s. **What import-time state now means.** Anything computed before `app.run()` is shared by every session for the life of the version. The code was audited against this on the move, and nothing an admin can change without redeploying is read at import: - Skills, prompts, grants and settings are read from DynamoDB per request. - Gateway `list_tools` runs per request. - AgentCard and JWKS caches use wall-clock buckets and first fill on a request. - `__warmup__` arrives on `/invocations`, so it runs after restore. Keep it that way. A value that must differ per session (ids, timestamps, credentials) or that an admin can change belongs in the handler. A Gateway tool catalog fetched at import would freeze until the next deploy. **One platform-level effect, measured and harmless here.** ADOT's X-Ray id generator draws from Python's `random`, and every restored VM replays the same sequence. In the first 20 V2 sessions the same span id appeared in 18 different traces. Trace ids are unaffected: the platform passes trace context in (every root `POST /invocations` span has a `parentSpanId`), so no trace id and no `(traceId, spanId)` pair was shared between sessions, and nothing in this repo joins telemetry on the span id alone. A trace started locally, without the incoming context (for example, on a thread that does not inherit it), would get colliding trace ids. **A race V2 made visible.** The chatbot's login `__warmup__` and the side panel's Files listing (`InvokeAgentRuntimeCommand`) hit the same brand-new session id at the same moment. On V2 the loser gets `409 RetryableConflictException: Session operation in progress, please retry` immediately. On V1 the same race returned 424 after about 5.5s, 4 times out of 4, so the warmup was already being lost. The refusal comes before the container is entered, so both callers now repeat the call: `doSignedPost` (warmup and chat turns) and the Files listing. They go through `chatbot/src/api/sessionConflict.ts`, which waits 0.5, 1, 2, 4 and 8s between attempts. The budget has to outlast a session create, measured at up to ~12.6s, and the SDK's own retry of this error gives up in under a second. Only this one 409 is retried. Observed live after the change: a warmup that lost three times got `warmup_ok` on the fourth attempt about 6s after page load, and a Files listing that lost six times got through on the seventh. ### 9.3 Runtime Configuration Injection A key design challenge: React apps need environment-specific values (API endpoints, Cognito IDs) that are only known after CDK deploys the resources. The solution: 1. **Build time**: Webpack bundles the React app. `config.ts` reads from `window.__CONFIG__` 2. **Deploy time**: CDK custom resource writes `config.js` to S3 with actual values 3. **Runtime**: `index.html` loads `