Skip to content

Repository files navigation

Track and Optimize Generative AI Spend on AWS

Workshop Cover

Sample code for the AWS Workshop: Track and Optimize Generative AI Spend on AWS

Level: 300 – Advanced | Duration: 3 hours

Overview

As organizations scale their generative AI workloads on Amazon Bedrock, a common challenge emerges: who is spending what, and where? Whether you're running multi-tenant applications, enabling developers with Claude Code, or deploying autonomous agents with Amazon Bedrock AgentCore, you need visibility into how your AI budget is being consumed.

This workshop is organized into two parts. Track provides hands-on experience with six AWS-native cost attribution mechanisms plus LiteLLM as a third-party option, so you know who is spending what, and where. Optimize then shows how to reduce that spend using Amazon Bedrock's built-in cost-reduction levers, without sacrificing quality. You'll learn to implement each mechanism, combine them for real-world scenarios, and build cost dashboards for operational visibility.

The Track samples live under samples/1-track/ and the Optimize samples under samples/2-optimize/, the latter organized into low-, medium-, and high-effort tiers.

For a detailed, hands-on walkthrough of the concepts covered here, see the companion workshop: Track and Optimize Generative AI Spend on AWS.

The Challenge

Generative AI spend is uniquely difficult to track:

  • Token-based pricing makes costs unpredictable compared to fixed-resource services
  • Shared infrastructure (one account, many teams) obscures who is driving costs
  • Agentic workloads (Claude Code, AgentCore) make many inference calls per task with no built-in association between business context and the API call
  • Multi-provider environments add complexity when teams use Bedrock alongside other model providers

Cost Attribution Mechanisms Covered

Mechanism Endpoint Visibility Latency Cost Type Granularity
IAM Principal Attribution bedrock-runtime, bedrock-mantle Up to 24h Billed dollars Per identity, per day
Application Inference Profiles bedrock-runtime Up to 24h Billed dollars Per profile, per day
Workspaces bedrock-mantle Up to 24h Billed dollars Per workspace, per day
Projects bedrock-mantle Up to 24h Billed dollars Per project, per day
Per-Request Metadata Tagging bedrock-runtime Near real-time Token counts Per request
IAM Identity Log Attribution bedrock-runtime Near real-time Token counts Per identity
LiteLLM (third-party) Proxy layer Real-time Estimated cost Per request

Optimization Levers Covered

Once you can see where spend goes, the optimize samples show how to reduce it, organized by how much work each lever takes to adopt:

Tier Levers
Low effort Model selection, prompt design, parameter tuning, prompt caching, adaptive thinking
Medium effort LLM routing, Bedrock Guardrails, RAG / indexing, batch inference
High effort Harness engineering, sub-agent delegation

Repository Structure

cfn/
└── workshop-stack.yaml                   # CloudFormation template for the workshop environment
samples/
├── 1-track/                              # Cost attribution: who is spending what, and where
│   ├── 1-iam-principal-attribution/      # IAM tagging and per-developer cost tracking
│   ├── 2-application-inference-profiles/ # Profile creation and traffic routing
│   ├── 3-workspaces/                     # Workspaces for Anthropic Messages API
│   ├── 4-projects/                       # Projects for OpenAI Responses API
│   ├── 5-per-request-metadata-tagging/   # Per-request metadata and log queries
│   ├── 6-iam-identity-log-attribution/   # Model invocation logging with IAM caller identity
│   └── 7-litellm/                        # LiteLLM proxy for multi-provider tracking
└── 2-optimize/                           # Cost reduction: spend less without sacrificing quality
    ├── 01-low-effort/                    # Model selection, prompt design, tuning, caching, adaptive thinking
    ├── 02-medium-effort/                 # LLM routing, guardrails, RAG/indexing, batch inference
    └── 03-high-effort/                   # Harness engineering, sub-agent delegation

Prerequisites

  • AWS account with Amazon Bedrock model access enabled
  • Basic familiarity with IAM roles and policies
  • Basic understanding of Amazon Bedrock APIs (Converse, InvokeModel)
  • Python 3.12+ and AWS CLI v2 configured

Getting Started

  1. Clone this repository:
git clone https://github.com/aws-samples/sample-code-for-genai-cost-management-workshop.git
cd sample-code-for-genai-cost-management-workshop
  1. Create a virtual environment and install dependencies:
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
  1. Configure your environment variables:
export AWS_REGION="us-east-1"
  1. Deploy the workshop infrastructure:

    Follow the instructions in cfn/README.md to deploy the CloudFormation stack.

  2. Work through the samples/1-track/ samples to learn attribution, then the samples/2-optimize/ samples to reduce spend - sequentially, or jump to any method independently

Learning Objectives

By the end of this workshop, you will be able to:

  1. Explain the cost attribution mechanisms available in Amazon Bedrock and when to use each
  2. Configure IAM principal attribution to track per-developer spend (including Claude Code users)
  3. Create and tag application inference profiles for per-application cost allocation
  4. Set up workspaces and projects for bedrock-mantle workloads
  5. Implement per-request metadata tagging for fine-grained per-prompt tracking
  6. Attribute costs for Amazon AgentCore agent workloads across tasks and sessions
  7. Combine multiple attribution methods for complete cost visibility
  8. Build cost dashboards using Cost Explorer, CUR 2.0, and QuickSight
  9. Reduce spend with Bedrock cost-optimization levers - model selection, prompt design and caching, routing, guardrails, RAG, batch inference, and agent-harness patterns

Target Audience

  • Cloud architects and platform engineers responsible for AI cost governance
  • FinOps practitioners managing generative AI budgets
  • DevOps and ML engineers deploying Bedrock-based applications, Claude Code, or AgentCore agents
  • Technical account managers and solutions architects advising on cost optimization

Security

See CONTRIBUTING for more information.

License

This library is licensed under the MIT-0 License. See the LICENSE file.

About

Track and Optimize Generative AI Spend on AWS Workshop

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors