Running LLM Inference on AWS: GPU Autoscaling, RAG, and Bedrock

Self-hosted GPU inference or managed models via Bedrock? An architecture view of serving LLMs in production, including RAG with vector stores and keeping costs sane.

Running LLM Inference on AWS: GPU Autoscaling, RAG, and Bedrock
Tomislav Pree
June 30, 2026
/
Articles

Managed vs self-hosted

The first architecture decision I make on any LLM project is whether to reach for a managed service or run inference myself, not which model to use. Bedrock gives me managed access to a range of foundation models without owning any GPU capacity, patching, or scaling logic, and for most product use cases that's the right default: I don't want to operate model-serving infrastructure if a managed API meets the latency and control requirements. Where I move to containerized GPU services instead is when I need something Bedrock doesn't offer: a specific fine-tuned or open-weight model, tighter control over batching and latency, or data residency requirements that rule out a shared managed endpoint. That's a real infrastructure commitment, so I only take it on when the managed path genuinely can't satisfy the requirement.

Autoscaling GPU workloads

Self-hosted GPU inference has to scale on the right signal, and CPU utilization is the wrong one. A GPU-bound model can sit at low CPU while the actual bottleneck, the GPU itself, is saturated. I scale containerized inference services on queue depth and request latency instead, so the system reacts to the thing that's actually degrading. GPU instances also have real cold-start cost: loading model weights onto a fresh instance takes meaningfully longer than a typical container start, so I keep a warm minimum capacity sized to expected baseline traffic rather than scaling from zero, and let autoscaling handle the burst above that floor.

RAG as an architecture pattern

Retrieval-augmented generation is where I spend most of my design time, because it's less a feature than an architecture decision about where knowledge lives. Rather than relying on what a model learned during training, a RAG pipeline retrieves relevant context from a vector store at request time and feeds it into the prompt alongside the user's query. The retrieval step sits in the request path before generation (embed the query, search the vector store for the nearest matches, assemble that context into the prompt), which means retrieval latency and quality directly bound the quality of the final answer. Getting the chunking and embedding strategy right upstream matters more than any prompt tuning downstream of it.

Cost control from day one

GPU capacity is expensive enough that cost has to be an architecture concern from the start, not a cleanup pass later. I right-size instance types to the model's actual memory and throughput needs rather than defaulting to the largest available GPU, and I route interruption-tolerant work (batch embedding jobs, offline evaluation) onto spot capacity, keeping only latency-sensitive live inference on on-demand instances. Observability has to cover tokens and GPU utilization specifically, not just request count and response time, because those are the two numbers that actually predict your bill and tell you whether capacity is sized correctly.

Takeaways

  • Default to managed model access via Bedrock. Move to self-hosted GPU serving only when you need control Bedrock doesn't give you.
  • Scale GPU inference on queue depth and latency, not CPU, and keep a warm floor to absorb cold starts.
  • RAG is an architecture pattern. Retrieval quality upstream bounds generation quality downstream.
  • Track token throughput and GPU utilization explicitly, and use spot capacity for interruption-tolerant batch work.