Sample code for the AWS Workshop: Track and Optimize Generative AI Spend on AWS
Level: 300 – Advanced | Duration: 3 hours
As organizations scale their generative AI workloads on Amazon Bedrock, a common challenge emerges: who is spending what, and where? Whether you're running multi-tenant applications, enabling developers with Claude Code, or deploying autonomous agents with Amazon Bedrock AgentCore, you need visibility into how your AI budget is being consumed.
This workshop is organized into two parts. Track provides hands-on experience with six AWS-native cost attribution mechanisms plus LiteLLM as a third-party option, so you know who is spending what, and where. Optimize then shows how to reduce that spend using Amazon Bedrock's built-in cost-reduction levers, without sacrificing quality. You'll learn to implement each mechanism, combine them for real-world scenarios, and build cost dashboards for operational visibility.
The Track samples live under
samples/1-track/and the Optimize samples undersamples/2-optimize/, the latter organized into low-, medium-, and high-effort tiers.
For a detailed, hands-on walkthrough of the concepts covered here, see the companion workshop: Track and Optimize Generative AI Spend on AWS.
Generative AI spend is uniquely difficult to track:
- Token-based pricing makes costs unpredictable compared to fixed-resource services
- Shared infrastructure (one account, many teams) obscures who is driving costs
- Agentic workloads (Claude Code, AgentCore) make many inference calls per task with no built-in association between business context and the API call
- Multi-provider environments add complexity when teams use Bedrock alongside other model providers
| Mechanism | Endpoint | Visibility Latency | Cost Type | Granularity |
|---|---|---|---|---|
| IAM Principal Attribution | bedrock-runtime, bedrock-mantle | Up to 24h | Billed dollars | Per identity, per day |
| Application Inference Profiles | bedrock-runtime | Up to 24h | Billed dollars | Per profile, per day |
| Workspaces | bedrock-mantle | Up to 24h | Billed dollars | Per workspace, per day |
| Projects | bedrock-mantle | Up to 24h | Billed dollars | Per project, per day |
| Per-Request Metadata Tagging | bedrock-runtime | Near real-time | Token counts | Per request |
| IAM Identity Log Attribution | bedrock-runtime | Near real-time | Token counts | Per identity |
| LiteLLM (third-party) | Proxy layer | Real-time | Estimated cost | Per request |
Once you can see where spend goes, the optimize samples show how to reduce it, organized by how much work each lever takes to adopt:
| Tier | Levers |
|---|---|
| Low effort | Model selection, prompt design, parameter tuning, prompt caching, adaptive thinking |
| Medium effort | LLM routing, Bedrock Guardrails, RAG / indexing, batch inference |
| High effort | Harness engineering, sub-agent delegation |
cfn/
└── workshop-stack.yaml # CloudFormation template for the workshop environment
samples/
├── 1-track/ # Cost attribution: who is spending what, and where
│ ├── 1-iam-principal-attribution/ # IAM tagging and per-developer cost tracking
│ ├── 2-application-inference-profiles/ # Profile creation and traffic routing
│ ├── 3-workspaces/ # Workspaces for Anthropic Messages API
│ ├── 4-projects/ # Projects for OpenAI Responses API
│ ├── 5-per-request-metadata-tagging/ # Per-request metadata and log queries
│ ├── 6-iam-identity-log-attribution/ # Model invocation logging with IAM caller identity
│ └── 7-litellm/ # LiteLLM proxy for multi-provider tracking
└── 2-optimize/ # Cost reduction: spend less without sacrificing quality
├── 01-low-effort/ # Model selection, prompt design, tuning, caching, adaptive thinking
├── 02-medium-effort/ # LLM routing, guardrails, RAG/indexing, batch inference
└── 03-high-effort/ # Harness engineering, sub-agent delegation
- AWS account with Amazon Bedrock model access enabled
- Basic familiarity with IAM roles and policies
- Basic understanding of Amazon Bedrock APIs (Converse, InvokeModel)
- Python 3.12+ and AWS CLI v2 configured
- Clone this repository:
git clone https://github.com/aws-samples/sample-code-for-genai-cost-management-workshop.git
cd sample-code-for-genai-cost-management-workshop- Create a virtual environment and install dependencies:
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt- Configure your environment variables:
export AWS_REGION="us-east-1"-
Deploy the workshop infrastructure:
Follow the instructions in
cfn/README.mdto deploy the CloudFormation stack. -
Work through the
samples/1-track/samples to learn attribution, then thesamples/2-optimize/samples to reduce spend - sequentially, or jump to any method independently
By the end of this workshop, you will be able to:
- Explain the cost attribution mechanisms available in Amazon Bedrock and when to use each
- Configure IAM principal attribution to track per-developer spend (including Claude Code users)
- Create and tag application inference profiles for per-application cost allocation
- Set up workspaces and projects for bedrock-mantle workloads
- Implement per-request metadata tagging for fine-grained per-prompt tracking
- Attribute costs for Amazon AgentCore agent workloads across tasks and sessions
- Combine multiple attribution methods for complete cost visibility
- Build cost dashboards using Cost Explorer, CUR 2.0, and QuickSight
- Reduce spend with Bedrock cost-optimization levers - model selection, prompt design and caching, routing, guardrails, RAG, batch inference, and agent-harness patterns
- Cloud architects and platform engineers responsible for AI cost governance
- FinOps practitioners managing generative AI budgets
- DevOps and ML engineers deploying Bedrock-based applications, Claude Code, or AgentCore agents
- Technical account managers and solutions architects advising on cost optimization
See CONTRIBUTING for more information.
This library is licensed under the MIT-0 License. See the LICENSE file.
