Single AI agents are powerful. Multi-agent systems are transformative. This guide builds a complete, production-ready multi-agent AI system on Amazon EKS, the exact architecture used by financial services, healthcare, and enterprise SaaS teams in 2026. Five specialised agents (Customer, Payment, Fraud, Compliance, Customer Support), a Custom Agent Router, Bedrock Guardrails, RAG with OpenSearch, response verification, WAF protection, and human approval gates. Every component configured from scratch. Every line of code included.What You Are Building Look at the architecture diagram. Every component in it has a job, and they all work together to serve a user request safely, accurately, and at enterprise scale. Here is the full data flow, before a single line of code: User → App (mobile/web) → Route 53 (DNS lookup) → CloudFront (CDN, content closer to user) → WAF (blocks malicious traffic) → Cognito (authenticates user, issues token) → API Gateway (receives task + login token) → Private VPC Link (secure internal routing) → ALB (distributes to healthy agent pods) → EKS (runs the containerised agent application) → Custom Agent Router (chooses the right agent) → Specialist Agent (Customer / Payment / Fraud / Compliance / Support) → Company Knowledge (S3 → OpenSearch → Bedrock Knowledge Bases) → Model & Safety (Bedrock Guardrails → Claude via Bedrock) → Business Tools (AgentCore Gateway → Business APIs) → Human Approval (for sensitive actions — Set Functions → Human) → Response Verification (custom code on EKS checks the response) → If rejected: return to agent for correction → If accepted: send to user → User sees the response That is the system. This guide builds every layer of it, step by step. What you need: AWS account with admin access AWS CLI configured (aws configure) kubectl installed Python 3.11+ Docker About 3–4 hours AWS services used: Route 53, CloudFront, WAF, S3, Cognito, API Gateway, VPC, ALB, EKS, Amazon Bedrock, Bedrock Knowledge Bases, OpenSearch Serverless, Bedrock Guardrails, AgentCore Gateway, Lambda, Step Functions, CloudWatch, IAM Step 1: Foundation VPC, EKS Cluster, and IAM Everything runs inside a VPC. Start here. 1.1 Create the VPC bash REGION="eu-west-1" ACCOUNT_ID=$(aws sts get-caller-identity --query Account --output text) CLUSTER_NAME="multi-agent-prod" # Create VPC with public and private subnets aws ec2 create-vpc \ --cidr-block 10.0.0.0/16 \ --tag-specifications 'ResourceType=vpc,Tags=[{Key=Name,Value=multi-agent-vpc},{Key=Project,Value=multi-agent}]' \ --region $REGION VPC_ID=$(aws ec2 describe-vpcs \ --filters "Name=tag:Name,Values=multi-agent-vpc" \ --query 'Vpcs[0].VpcId' \ --output text \ --region $REGION) echo "VPC ID: $VPC_ID" # Enable DNS hostnames (required for EKS) aws ec2 modify-vpc-attribute \ --vpc-id $VPC_ID \ --enable-dns-hostnames \ --region $REGION # Create Internet Gateway for public subnets IGW_ID=$(aws ec2 create-internet-gateway \ --query 'InternetGateway.InternetGatewayId' \ --output text \ --region $REGION) aws ec2 attach-internet-gateway \ --internet-gateway-id $IGW_ID \ --vpc-id $VPC_ID \ --region $REGION # Create public subnets (for ALB and NAT Gateway) PUBLIC_SUBNET_1=$(aws ec2 create-subnet \ --vpc-id $VPC_ID \ --cidr-block 10.0.1.0/24 \ --availability-zone ${REGION}a \ --tag-specifications 'ResourceType=subnet,Tags=[{Key=Name,Value=public-1a},{Key=kubernetes.io/role/elb,Value=1}]' \ --query 'Subnet.SubnetId' \ --output text \ --region $REGION) PUBLIC_SUBNET_2=$(aws ec2 create-subnet \ --vpc-id $VPC_ID \ --cidr-block 10.0.2.0/24 \ --availability-zone ${REGION}b \ --tag-specifications 'ResourceType=subnet,Tags=[{Key=Name,Value=public-1b},{Key=kubernetes.io/role/elb,Value=1}]' \ --query 'Subnet.SubnetId' \ --output text \ --region $REGION) # Create private subnets (for EKS nodes — never directly exposed) PRIVATE_SUBNET_1=$(aws ec2 create-subnet \ --vpc-id $VPC_ID \ --cidr-block 10.0.10.0/24 \ --availability-zone ${REGION}a \ --tag-specifications 'ResourceType=subnet,Tags=[{Key=Name,Value=private-1a},{Key=kubernetes.io/role/internal-elb,Value=1}]' \ --query 'Subnet.SubnetId' \ --output text \ --region $REGION) PRIVATE_SUBNET_2=$(aws ec2 create-subnet \ --vpc-id $VPC_ID \ --cidr-block 10.0.11.0/24 \ --availability-zone ${REGION}b \ --tag-specifications 'ResourceType=subnet,Tags=[{Key=Name,Value=private-1b},{Key=kubernetes.io/role/internal-elb,Value=1}]' \ --query 'Subnet.SubnetId' \ --output text \ --region $REGION) echo "Subnets created: $PUBLIC_SUBNET_1, $PUBLIC_SUBNET_2, $PRIVATE_SUBNET_1, $PRIVATE_SUBNET_2" 1.2 Create the EKS Cluster bash # Create the EKS cluster using eksctl (simplest approach) # Install eksctl if not present: # brew install eksctl (macOS) # or download from https://eksctl.io cat > cluster-config.yaml agent-pod-trust.json agent-pod-policy.json company-docs/policies/refund-policy.txt company-docs/procedures/account-verification.txt company-docs/products/subscription-tiers.txt opensearch-policy.json bedrock-kb-trust.json tuple[str, list]: """ Retrieve relevant company knowledge from Bedrock Knowledge Base. This is the RAG step — find relevant documents before calling the model. """ try: response = self.bedrock_agent.retrieve( knowledgeBaseId=KNOWLEDGE_BASE_ID, retrievalQuery={"text": query}, retrievalConfiguration={ "vectorSearchConfiguration": { "numberOfResults": num_results, "overrideSearchType": "HYBRID" } } ) results = response.get("retrievalResults", []) if not results: return "", [] # Format retrieved context context_parts = [] sources = [] for i, result in enumerate(results, 1): content = result["content"]["text"] source = result.get("location", {}).get("s3Location", {}).get("uri", "company-docs") score = result.get("score", 0) context_parts.append(f"[Source {i}: {source} (relevance: {score:.2f})]\n{content}") sources.append({"source": source, "score": score}) return "\n\n".join(context_parts), sources except Exception as e: logger.error(f"Knowledge retrieval failed: {e}") return "", [] def call_bedrock_with_guardrails( self, messages: list, additional_context: str = "", max_tokens: int = 1000 ) -> tuple[str, int]: """ Call Claude via Bedrock with Guardrails applied. Guardrails check both the input and the output. """ full_system = self.system_prompt if additional_context: full_system += f"\n\n## Relevant Company Knowledge:\n{additional_context}" body = { "anthropic_version": "bedrock-2023-05-31", "max_tokens": max_tokens, "system": full_system, "messages": messages } # Apply guardrails to the request if GUARDRAIL_ID: body["amazon-bedrock-guardrailConfig"] = { "guardrailIdentifier": GUARDRAIL_ID, "guardrailVersion": GUARDRAIL_VERSION, "trace": "enabled" } start_time = time.time() response = self.bedrock.invoke_model( modelId=MODEL_ID, body=json.dumps(body) ) latency_ms = int((time.time() - start_time) * 1000) result = json.loads(response['body'].read()) # Check if guardrails blocked the response guardrail_action = result.get("amazon-bedrock-guardrailAction", "NONE") if guardrail_action == "GUARDRAIL_INTERVENED": logger.warning(f"Guardrail intervened for {self.agent_type}") content = result.get("content", [{}])[0].get("text", "I cannot provide that response due to our content policy.") else: content = result["content"][0]["text"] total_tokens = result["usage"]["input_tokens"] + result["usage"]["output_tokens"] return content, total_tokens def check_requires_human_approval(self, request: str, response: str) -> tuple[bool, str]: """ Determine if this response requires human approval before being sent. Each specialist agent overrides this with its own rules. """ return False, "" def handle(self, request: str, session_id: str, user_id: str) -> AgentResponse: """Main handler — called by the Agent Router""" start_time = time.time() logger.info(f"{self.agent_type} handling request for session {session_id}") # Step 1: Retrieve relevant company knowledge (RAG) knowledge_context, sources = self.retrieve_company_knowledge(request) # Step 2: Build conversation messages messages = [{"role": "user", "content": f"User ID: {user_id}\n\nRequest: {request}"}] # Step 3: Call the model with guardrails content, tokens_used = self.call_bedrock_with_guardrails( messages=messages, additional_context=knowledge_context ) # Step 4: Check human approval requirement requires_approval, approval_reason = self.check_requires_human_approval(request, content) latency_ms = int((time.time() - start_time) * 1000) # Step 5: Publish metrics self.publish_metrics(tokens_used, latency_ms, requires_approval) return AgentResponse( agent_type=self.agent_type, success=True, content=content, requires_human_approval=requires_approval, approval_reason=approval_reason, knowledge_sources=sources, tokens_used=tokens_used, latency_ms=latency_ms, session_id=session_id ) def publish_metrics(self, tokens: int, latency: int, requires_approval: bool): """Publish CloudWatch metrics for cost and performance monitoring""" try: self.cloudwatch.put_metric_data( Namespace="MultiAgentSystem/Agents", MetricData=[ { "MetricName": "TokensConsumed", "Value": tokens, "Unit": "Count", "Dimensions": [{"Name": "AgentType", "Value": self.agent_type}] }, { "MetricName": "RequestLatencyMs", "Value": latency, "Unit": "Milliseconds", "Dimensions": [{"Name": "AgentType", "Value": self.agent_type}] }, { "MetricName": "HumanApprovalRequired", "Value": 1 if requires_approval else 0, "Unit": "Count", "Dimensions": [{"Name": "AgentType", "Value": self.agent_type}] } ] ) except Exception as e: logger.warning(f"Metrics publish failed: {e}") python # agents/specialist_agents.py # The five specialist agents from your architecture diagram from .base_agent import BaseAgent import re class CustomerAgent(BaseAgent): """ Handles customer account requests: - Account information queries - Profile updates - Subscription changes - Billing history """ agent_type = "customer" system_prompt = """You are the Customer Account Agent for our platform. Your responsibilities: - Answer questions about the customer's account, subscription, and billing history - Help customers update their profile information - Explain subscription tiers and help customers choose the right plan - Process subscription upgrades and downgrades per company policy Always: - Verify you are speaking about the correct account (reference User ID) - Be accurate about subscription terms and pricing from company documentation - Never process payments directly — route to Payment Agent - Never access other customers' information You have access to company knowledge about subscription tiers, policies, and procedures.""" def check_requires_human_approval(self, request: str, response: str) -> tuple[bool, str]: """Account changes above certain thresholds need approval""" # Account deletion always needs human confirmation if any(term in request.lower() for term in ["delete account", "close account", "cancel everything"]): return True, "Account deletion requires human confirmation and data retention review" return False, "" class PaymentAgent(BaseAgent): """ Handles approved payment tasks: - Refund processing - Payment dispute investigation - Invoice generation - Subscription billing issues """ agent_type = "payment" system_prompt = """You are the Payment Agent for our platform. Your responsibilities: - Process refund requests according to company refund policy - Investigate payment disputes - Clarify billing questions and invoice details - Process approved payment tasks Rules you MUST follow: - Refunds up to £500: you can approve and process directly - Refunds above £500: flag for human manager approval with full justification - Duplicate charges: always refund immediately, no approval needed - Suspected fraud: immediately escalate to Fraud Agent — do NOT process payment - Always reference the specific transaction ID in your response - Log all payment actions with timestamp and reason Company refund policy is available in your knowledge base.""" def check_requires_human_approval(self, request: str, response: str) -> tuple[bool, str]: """Large refunds and unusual payment patterns need human review""" # Extract amounts from request and response amounts = re.findall(r'£(\d+(?:,\d{3})*(?:\.\d{2})?)', request + response) for amount_str in amounts: amount = float(amount_str.replace(',', '')) if amount > self.high_value_threshold: return True, f"Refund of £{amount:,.2f} exceeds £{self.high_value_threshold:.0f} threshold — manager approval required" return False, "" class FraudAgent(BaseAgent): """ Investigates suspicious activity: - Unusual transaction patterns - Suspicious login attempts - KYC/AML checks - Sanctions screening """ agent_type = "fraud" system_prompt = """You are the Fraud Investigation Agent for our platform. Your responsibilities: - Investigate reports of suspicious account activity - Identify patterns consistent with fraud, money laundering, or identity theft - Check accounts against KYC (Know Your Customer) requirements - Flag potential AML (Anti-Money Laundering) concerns - Perform sanctions and PEP (Politically Exposed Person) screening Investigation framework: 1. Review the transaction pattern or suspicious activity reported 2. Check against company AML and KYC policies (in knowledge base) 3. Assign a risk level: LOW / MEDIUM / HIGH / CRITICAL 4. For MEDIUM and above: recommend specific actions 5. For HIGH and CRITICAL: ALWAYS escalate for human review You do NOT have authority to freeze accounts or block transactions. You gather evidence and make recommendations. Human team executes. Your knowledge base contains AML procedures, KYC requirements, and compliance documentation.""" def check_requires_human_approval(self, request: str, response: str) -> tuple[bool, str]: """Fraud findings always need human review before action""" risk_indicators = ["HIGH", "CRITICAL", "suspicious", "fraud", "money laundering", "sanctions", "freeze", "block", "escalate"] response_lower = response.lower() for indicator in risk_indicators: if indicator.lower() in response_lower: return True, f"Fraud investigation flagged risk indicator: '{indicator}' — compliance team review required" return False, "" class ComplianceAgent(BaseAgent): """ Handles regulatory compliance checks: - KYC and AML verification - Sanctions screening - Regulatory reporting requirements - Data protection queries (GDPR, etc.) """ agent_type = "compliance" system_prompt = """You are the Compliance Agent for our platform. Your responsibilities: - Verify customer compliance with KYC and AML requirements - Assess regulatory obligations for specific transactions or business activities - Advise on data protection requirements (GDPR, UK GDPR) - Check against sanctions lists and PEP databases - Advise on regulatory reporting thresholds Important boundaries: - You provide compliance INFORMATION and ASSESSMENT, not legal advice - For definitive legal interpretations, recommend customers consult a qualified lawyer - High-risk compliance findings MUST be reviewed by the compliance team (human) - Always cite the specific regulation or company policy you are referencing Your knowledge base contains regulatory guidelines, compliance procedures, and policy documents.""" def check_requires_human_approval(self, request: str, response: str) -> tuple[bool, str]: """Compliance actions always require human sign-off""" compliance_actions = ["report to", "file with", "notify regulator", "submit to fca", "suspicious activity report", "sar", "freeze", "restrict"] response_lower = response.lower() for action in compliance_actions: if action in response_lower: return True, f"Compliance action '{action}' requires regulatory compliance team approval" return False, "" class CustomerSupportAgent(BaseAgent): """ Handles general customer support requests: - Product questions - Technical troubleshooting - Feature guidance - Escalation routing """ agent_type = "customer_support" system_prompt = """You are the Customer Support Agent for our platform. Your responsibilities: - Answer product and feature questions - Guide customers through technical issues step by step - Explain how to use platform features effectively - Identify when issues need escalation to specialist agents Routing rules: - Billing and payment questions → route to Payment Agent - Account access problems → route to Customer Agent - Suspicious activity → route to Fraud Agent - Compliance questions → route to Compliance Agent - General product help → handle yourself Response guidelines: - Be friendly, clear, and patient - Use numbered steps for technical instructions - Always confirm the customer understood the solution - If you cannot resolve in 3 exchanges, offer to escalate Your knowledge base contains product documentation, FAQs, and troubleshooting guides.""" def check_requires_human_approval(self, request: str, response: str) -> tuple[bool, str]: """Escalation requests need human agent involvement""" escalation_terms = ["speak to a human", "talk to someone", "escalate", "manager", "supervisor", "not helpful", "frustrated"] request_lower = request.lower() for term in escalation_terms: if term in request_lower: return True, "Customer has requested human agent involvement" return False, "" Step 5: The Custom Agent Router The blue brain icon in the centre of your diagram. This is what receives every request and decides which specialist agent handles it. python # agents/router.py # The Custom Agent Router — decides which agent handles each request import boto3 import json import logging import os from .specialist_agents import ( CustomerAgent, PaymentAgent, FraudAgent, ComplianceAgent, CustomerSupportAgent ) from .base_agent import AgentResponse, bedrock logger = logging.getLogger(__name__) MODEL_ID = "anthropic.claude-3-5-haiku-20241022" # Fast model for routing decisions class AgentRouter: """ The Custom Agent Router coordinates the specialist agents. Responsibilities: 1. Receive user requests from the API 2. Classify the request and select the right agent 3. Dispatch to the selected agent 4. Receive the agent's response 5. Pass the response to Response Verification 6. Handle human approval workflows """ def __init__(self): self.agents = { "customer": CustomerAgent(), "payment": PaymentAgent(), "fraud": FraudAgent(), "compliance": ComplianceAgent(), "customer_support": CustomerSupportAgent() } # Use a fast, cheap model for routing — no need for Opus here self.bedrock = bedrock def classify_request(self, request: str, user_id: str) -> str: """ Classify which specialist agent should handle this request. Uses Claude Haiku for speed — routing is a simple classification task. """ response = self.bedrock.invoke_model( modelId=MODEL_ID, body=json.dumps({ "anthropic_version": "bedrock-2023-05-31", "max_tokens": 50, "temperature": 0.0, # No randomness — routing must be deterministic "system": """Classify this customer request into exactly ONE category. Return ONLY the category name, nothing else. Categories: - customer: Account info, profile updates, subscription changes, login help - payment: Refunds, billing disputes, invoice questions, payment processing - fraud: Suspicious activity, unusual transactions, security concerns, identity theft - compliance: KYC, AML, regulatory questions, data protection, GDPR - customer_support: General product help, technical issues, feature questions, anything else If multiple categories apply, pick the PRIMARY one.""", "messages": [{"role": "user", "content": f"Request: {request}"}] }) ) result = json.loads(response['body'].read()) agent_type = result["content"][0]["text"].strip().lower() # Validate the classification if agent_type not in self.agents: logger.warning(f"Unknown agent type '{agent_type}' — defaulting to customer_support") return "customer_support" logger.info(f"Request classified as: {agent_type}") return agent_type def route(self, request: str, session_id: str, user_id: str) -> AgentResponse: """ Main routing method. Classifies the request and dispatches to the correct specialist agent. """ logger.info(f"Router received request for session {session_id}") # Step 1: Classify the request agent_type = self.classify_request(request, user_id) # Step 2: Dispatch to the selected agent selected_agent = self.agents[agent_type] response = selected_agent.handle(request, session_id, user_id) return response Step 6: Response Verification The Safety Net The centre-right box in your diagram: "Response Verification Custom Code on Amazon EKS". This checks every response before it reaches the user. python # verification/response_verifier.py # Custom verification code running on EKS import re import json import logging from dataclasses import dataclass from typing import Optional logger = logging.getLogger(__name__) @dataclass class VerificationResult: accepted: bool rejection_reason: Optional[str] = None modified_content: Optional[str] = None class ResponseVerifier: """ Checks agent responses before they reach the user. This is a second safety layer — after Bedrock Guardrails, this code performs custom business-logic checks that are specific to the application. From the architecture: "Checks the proposed response before releasing it to the user" "Send rejection reason" back to agent if rejected "Agent-generated response accepted" if it passes """ # Patterns that indicate a response should be blocked BLOCKED_PATTERNS = [ r'\bpassword\b.{0,50}\bis\b', # Never reveal passwords r'\b4[0-9]{12}(?:[0-9]{3})?\b', # Credit card numbers r'\b[0-9]{3}-[0-9]{2}-[0-9]{4}\b', # SSN format r']*>', # XSS attempts in output r'SELECT .+ FROM .+ WHERE', # SQL in output (injection indicator) r'internal server error', # Exposing system errors to users r'stack trace', # Exposing technical details r'(api_key|secret_key)\s*[:=]', # Exposing credentials ] # Financial accuracy patterns — dollar amounts must be precise FINANCIAL_PATTERNS = [ r'approximately\s+£', # "Approximately £500" is not okay for financial responses r'around\s+£', r'roughly\s+£', ] # Patterns that require human review despite agent passing them ESCALATION_PATTERNS = [ r'legal action', r'regulatory authority', r'media', r'going public', r'formal complaint', ] def verify(self, response_content: str, agent_type: str, original_request: str) -> VerificationResult: """ Verify an agent response before sending it to the user. Returns: - VerificationResult with accepted=True if the response passes - VerificationResult with accepted=False and rejection_reason if blocked """ # Check 1: Blocked patterns — hard stop for pattern in self.BLOCKED_PATTERNS: if re.search(pattern, response_content, re.IGNORECASE): logger.warning(f"Response blocked: pattern '{pattern}' detected in {agent_type} response") return VerificationResult( accepted=False, rejection_reason=f"Response contains blocked content pattern. Rewrite without including {pattern}" ) # Check 2: Financial responses must be precise if agent_type == "payment": for pattern in self.FINANCIAL_PATTERNS: if re.search(pattern, response_content, re.IGNORECASE): return VerificationResult( accepted=False, rejection_reason="Financial responses must use exact amounts, not approximations. Rewrite with precise figures." ) # Check 3: Minimum quality — response must be substantive if len(response_content.strip()) str: """ Submit a response for human approval. Returns an approval_id that can be used to check status. """ approval_id = str(uuid.uuid4()) # Store the approval request in DynamoDB approval_table.put_item(Item={ "approvalId": approval_id, "sessionId": session_id, "userId": user_id, "agentType": agent_type, "originalRequest": original_request, "proposedResponse": proposed_response, "approvalReason": approval_reason, "status": "PENDING", "createdAt": datetime.now(timezone.utc).isoformat(), "ttl": int(datetime.now(timezone.utc).timestamp()) + 3600 # 1 hour TTL }) # Notify the human approval team via SNS sns.publish( TopicArn=self.APPROVAL_TOPIC_ARN, Subject=f"APPROVAL REQUIRED: {agent_type.upper()} Agent — {approval_reason[:50]}", Message=json.dumps({ "approval_id": approval_id, "approval_url": f"https://admin.yourcompany.com/approvals/{approval_id}", "session_id": session_id, "user_id": user_id, "agent_type": agent_type, "reason": approval_reason, "original_request": original_request[:500], "proposed_response": proposed_response[:500], "created_at": datetime.now(timezone.utc).isoformat() }, indent=2), MessageAttributes={ "agent_type": { "DataType": "String", "StringValue": agent_type }, "priority": { "DataType": "String", "StringValue": "HIGH" if agent_type in ["fraud", "compliance"] else "MEDIUM" } } ) logger.info(f"Human approval requested: {approval_id} for session {session_id}") return approval_id def check_approval_status(self, approval_id: str) -> dict: """Check the status of a pending approval""" response = approval_table.get_item(Key={"approvalId": approval_id}) return response.get("Item", {"status": "NOT_FOUND"}) def process_approval_decision(self, approval_id: str, decision: str, approver_id: str, notes: str = "") -> bool: """ Record a human approval decision. Called by the approval UI when a human makes a decision. decision: "APPROVED" or "REJECTED" """ if decision not in ["APPROVED", "REJECTED"]: raise ValueError(f"Invalid decision: {decision}. Must be APPROVED or REJECTED") approval_table.update_item( Key={"approvalId": approval_id}, UpdateExpression="SET #s = :status, approverId = :approver, notes = :notes, decidedAt = :decided", ExpressionAttributeNames={"#s": "status"}, ExpressionAttributeValues={ ":status": decision, ":approver": approver_id, ":notes": notes, ":decided": datetime.now(timezone.utc).isoformat() } ) logger.info(f"Approval {approval_id} {decision} by {approver_id}") return decision == "APPROVED" Step 8: The Main API Handler Bringing It All Together python # main.py # FastAPI application running on EKS — the entry point from API Gateway from fastapi import FastAPI, HTTPException, Header, Depends from pydantic import BaseModel from typing import Optional import logging import uuid import time import os from agents.router import AgentRouter from verification.response_verifier import ResponseVerifier from approval.human_approval import HumanApprovalWorkflow logging.basicConfig(level=logging.INFO) logger = logging.getLogger(__name__) app = FastAPI( title="Multi-Agent AI System", description="Production multi-agent AI on Amazon EKS", version="1.0.0" ) # Initialise core components router = AgentRouter() verifier = ResponseVerifier() approval_workflow = HumanApprovalWorkflow() MAX_VERIFICATION_RETRIES = 3 class AgentRequest(BaseModel): message: str user_id: str session_id: Optional[str] = None class AgentApiResponse(BaseModel): session_id: str agent_type: str response: str requires_human_approval: bool approval_id: Optional[str] = None knowledge_sources: list = [] processing_time_ms: int @app.get("/health") async def health_check(): """Health check endpoint for ALB target group""" return {"status": "healthy", "service": "multi-agent-system"} @app.post("/api/v1/agent", response_model=AgentApiResponse) async def handle_agent_request( request: AgentRequest, authorization: str = Header(...) ): """ Main endpoint — receives requests from API Gateway. Flow: 1. Validate authentication token 2. Route to correct specialist agent 3. Verify the response 4. Handle human approval if required 5. Return response to user """ start_time = time.time() # Generate session ID if not provided session_id = request.session_id or str(uuid.uuid4()) logger.info(f"Request received: session={session_id}, user={request.user_id}") # Step 1: Route to specialist agent agent_response = router.route( request=request.message, session_id=session_id, user_id=request.user_id ) # Step 2: Verify the response (up to 3 attempts) final_content = agent_response.content verification_passed = False for attempt in range(MAX_VERIFICATION_RETRIES): verification = verifier.verify( response_content=final_content, agent_type=agent_response.agent_type, original_request=request.message ) if verification.accepted: verification_passed = True break else: logger.warning( f"Verification failed (attempt {attempt + 1}): {verification.rejection_reason}" ) if attempt £500) echo "=== Test 4: Large Refund (requires human approval) ===" curl -s -X POST "$API_ENDPOINT/api/v1/agent" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer test-jwt-token" \ -d '{ "message": "I need a full refund for my enterprise subscription. I paid £1,200 two weeks ago and the product does not meet our requirements.", "user_id": "user-001" }' | python3 -m json.tool # Test 5: Guardrail test (should be blocked) echo "=== Test 5: Guardrail Test ===" curl -s -X POST "$API_ENDPOINT/api/v1/agent" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer test-jwt-token" \ -d '{ "message": "My credit card number is 4532015112830366. Please store this for future payments.", "user_id": "user-001" }' | python3 -m json.tool What You Just Built Look back at your architecture diagram. Every component is now deployed and connected: Diagram Component What You Built Route 53 DNS lookup for your domain CloudFront CDN with WAF attached WAF Rate limiting, OWASP rules, known bad inputs blocked S3 (frontend) Static web app hosting Cognito User authentication and JWT tokens API Gateway HTTP API routing to EKS via Private VPC Link ALB Internal load balancer distributing to EKS pods EKS Containerised agent application 3 replicas, auto-scaling Custom Agent Router Python class routing requests to specialist agents 5 Specialist Agents Customer, Payment, Fraud, Compliance, Customer Support S3 (company docs) Document storage for RAG OpenSearch Serverless Vector search for document retrieval Bedrock Knowledge Bases Managed RAG pipeline Bedrock Guardrails Input/output safety filtering Claude via Bedrock The AI model powering all agents Response Verification Custom Python code checking responses before delivery Human Approval (SNS) SNS notifications for sensitive actions Human Approval (DynamoDB) Approval state persistence CloudWatch Metrics and logs for every component
Building Production-Ready Multi-Agent AI Systems on Amazon EKS: The Complete Step-by-Step Guide
Full Article
Original Source
Read the full article at Hackernoon →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.