[{"data":1,"prerenderedAt":296},["ShallowReactive",2],{"blog-post-en-llm-inference-aws-gpu-rag-bedrock":3,"blog-latest-en-llm-inference-aws-gpu-rag-bedrock":79},{"id":4,"title":5,"body":6,"category":68,"date":69,"description":60,"extension":70,"meta":71,"navigation":72,"order":73,"path":74,"seo":75,"stem":76,"summary":77,"__hash__":78},"blog\u002Fen\u002Fblog\u002Fllm-inference-aws-gpu-rag-bedrock.md","Running LLM Inference on AWS: GPU Autoscaling, RAG, and Bedrock",{"type":7,"value":8,"toc":59},"minimark",[9,14,18,22,25,29,32,36,39,43],[10,11,13],"h2",{"id":12},"managed-vs-self-hosted","Managed vs self-hosted",[15,16,17],"p",{},"The first architecture decision I make on any LLM project is whether to reach for a managed service or run inference myself, not which model to use. Bedrock gives me managed access to a range of foundation models without owning any GPU capacity, patching, or scaling logic, and for most product use cases that's the right default: I don't want to operate model-serving infrastructure if a managed API meets the latency and control requirements. Where I move to containerized GPU services instead is when I need something Bedrock doesn't offer: a specific fine-tuned or open-weight model, tighter control over batching and latency, or data residency requirements that rule out a shared managed endpoint. That's a real infrastructure commitment, so I only take it on when the managed path genuinely can't satisfy the requirement.",[10,19,21],{"id":20},"autoscaling-gpu-workloads","Autoscaling GPU workloads",[15,23,24],{},"Self-hosted GPU inference has to scale on the right signal, and CPU utilization is the wrong one. A GPU-bound model can sit at low CPU while the actual bottleneck, the GPU itself, is saturated. I scale containerized inference services on queue depth and request latency instead, so the system reacts to the thing that's actually degrading. GPU instances also have real cold-start cost: loading model weights onto a fresh instance takes meaningfully longer than a typical container start, so I keep a warm minimum capacity sized to expected baseline traffic rather than scaling from zero, and let autoscaling handle the burst above that floor.",[10,26,28],{"id":27},"rag-as-an-architecture-pattern","RAG as an architecture pattern",[15,30,31],{},"Retrieval-augmented generation is where I spend most of my design time, because it's less a feature than an architecture decision about where knowledge lives. Rather than relying on what a model learned during training, a RAG pipeline retrieves relevant context from a vector store at request time and feeds it into the prompt alongside the user's query. The retrieval step sits in the request path before generation (embed the query, search the vector store for the nearest matches, assemble that context into the prompt), which means retrieval latency and quality directly bound the quality of the final answer. Getting the chunking and embedding strategy right upstream matters more than any prompt tuning downstream of it.",[10,33,35],{"id":34},"cost-control-from-day-one","Cost control from day one",[15,37,38],{},"GPU capacity is expensive enough that cost has to be an architecture concern from the start, not a cleanup pass later. I right-size instance types to the model's actual memory and throughput needs rather than defaulting to the largest available GPU, and I route interruption-tolerant work (batch embedding jobs, offline evaluation) onto spot capacity, keeping only latency-sensitive live inference on on-demand instances. Observability has to cover tokens and GPU utilization specifically, not just request count and response time, because those are the two numbers that actually predict your bill and tell you whether capacity is sized correctly.",[10,40,42],{"id":41},"takeaways","Takeaways",[44,45,46,50,53,56],"ul",{},[47,48,49],"li",{},"Default to managed model access via Bedrock. Move to self-hosted GPU serving only when you need control Bedrock doesn't give you.",[47,51,52],{},"Scale GPU inference on queue depth and latency, not CPU, and keep a warm floor to absorb cold starts.",[47,54,55],{},"RAG is an architecture pattern. Retrieval quality upstream bounds generation quality downstream.",[47,57,58],{},"Track token throughput and GPU utilization explicitly, and use spot capacity for interruption-tolerant batch work.",{"title":60,"searchDepth":61,"depth":61,"links":62},"",2,[63,64,65,66,67],{"id":12,"depth":61,"text":13},{"id":20,"depth":61,"text":21},{"id":27,"depth":61,"text":28},{"id":34,"depth":61,"text":35},{"id":41,"depth":61,"text":42},"Articles","June 30, 2026","md",{},true,4,"\u002Fen\u002Fblog\u002Fllm-inference-aws-gpu-rag-bedrock",{"title":5,"description":60},"en\u002Fblog\u002Fllm-inference-aws-gpu-rag-bedrock","Self-hosted GPU inference or managed models via Bedrock? An architecture view of serving LLMs in production, including RAG with vector stores and keeping costs sane.","esmc5T-lSxzmasuOI0QGEI2EFuOC-_G7Mt6WHPBDHrQ",[80,165,232],{"id":81,"title":82,"body":83,"category":156,"date":157,"description":60,"extension":70,"meta":158,"navigation":72,"order":159,"path":160,"seo":161,"stem":162,"summary":163,"__hash__":164},"blog\u002Fen\u002Fblog\u002Fgitops-argocd-helm-regulated-environments.md","GitOps with ArgoCD and Helm: Auditable Releases in Regulated Environments",{"type":7,"value":84,"toc":149},[85,89,101,105,108,112,119,123,130,132],[10,86,88],{"id":87},"why-push-based-deployments-fail-audits","Why push-based deployments fail audits",[15,90,91,92,96,97,100],{},"I still see teams deploying to Kubernetes by running ",[93,94,95],"code",{},"kubectl apply"," from a laptop, or by letting a CI pipeline push straight into a cluster on merge. Both work, until someone asks the question every regulated environment eventually asks: what was running in production on a given date, who approved it, and can you prove it. With push-based deployments, the honest answer is usually \"we're not entirely sure.\" The cluster's actual state and the record of intended state drift apart the moment a manual ",[93,98,99],{},"kubectl edit"," or an emergency hotfix bypasses the pipeline. Auditors don't accept \"check the deployment logs\" as a satisfying answer, and neither should I.",[10,102,104],{"id":103},"the-gitops-model","The GitOps model",[15,106,107],{},"The fix I reach for is GitOps: the desired state of every cluster lives in a Git repository, and a reconciliation controller (I use ArgoCD) continuously compares that declared state against what's actually running, and corrects any drift automatically. Helm charts template the Kubernetes manifests, so a single chart describes an application while environment-specific values files supply the differences between dev, staging, and production. Nothing reaches a cluster except through a change to Git, and every change to Git is a reviewed merge request. That single constraint, no direct cluster access ever, is what turns \"trust me\" into \"here's the commit.\"",[10,109,111],{"id":110},"multi-environment-promotion","Multi-environment promotion",[15,113,114,115,118],{},"With this structure, promoting a release between environments stops being a deploy script and becomes a Git operation. The same chart version moves from staging to production by merging a values-file change that bumps the image tag or config for that environment. Rollback is a ",[93,116,117],{},"git revert",", not a scramble to remember the previous Helm release number. Because ArgoCD reconciles continuously, the moment the revert lands, the cluster follows, with no separate rollback tooling and no manual intervention. This also means the promotion history is the commit history: I can point at a pull request and show exactly when a version reached production and who approved it.",[10,120,122],{"id":121},"security-stages-in-ci","Security stages in CI",[15,124,125,126,129],{},"GitOps handles what happens after a change is approved, but I still want confidence in what's being proposed. I run build, test, and security scanning as distinct stages in GitLab CI before an image or chart is even eligible for a values-file bump. Dependency scanning, container image scanning, and policy checks all gate the pipeline. Built artifacts are pushed to a central registry and referenced by immutable tag, never by ",[93,127,128],{},"latest",", so what's described in Git is exactly what gets deployed, with no ambiguity about which build is running.",[10,131,42],{"id":41},[44,133,134,137,140,143,146],{},[47,135,136],{},"Declarative desired state in Git beats imperative deployment scripts for anything that needs to be audited later.",[47,138,139],{},"No manual cluster access. Every change to running infrastructure is a reviewed, merged commit.",[47,141,142],{},"Promotion and rollback become Git operations: a merge or a revert, not bespoke tooling.",[47,144,145],{},"Security scanning belongs in CI, before a change is even proposed for deployment.",[47,147,148],{},"The audit trail falls out of how the deployment model works, so it is not a separate system to maintain.",{"title":60,"searchDepth":61,"depth":61,"links":150},[151,152,153,154,155],{"id":87,"depth":61,"text":88},{"id":103,"depth":61,"text":104},{"id":110,"depth":61,"text":111},{"id":121,"depth":61,"text":122},{"id":41,"depth":61,"text":42},"Tutorials","August 11, 2026",{},1,"\u002Fen\u002Fblog\u002Fgitops-argocd-helm-regulated-environments",{"title":82,"description":60},"en\u002Fblog\u002Fgitops-argocd-helm-regulated-environments","How declarative, Git-driven deployments make multi-environment releases reviewable, auditable, and reversible, and why regulated industries need exactly that.","ZZ2hInSupJyt7pIzUsNQPxVWs7OOfE3aK_E77XpXRU8",{"id":166,"title":167,"body":168,"category":68,"date":225,"description":60,"extension":70,"meta":226,"navigation":72,"order":61,"path":227,"seo":228,"stem":229,"summary":230,"__hash__":231},"blog\u002Fen\u002Fblog\u002Faws-multi-account-terraform.md","Designing AWS Multi-Account Architectures with Terraform",{"type":7,"value":169,"toc":218},[170,174,177,181,184,188,195,199,202,204],[10,171,173],{"id":172},"why-one-account-is-never-enough","Why one account is never enough",[15,175,176],{},"Early on, it's tempting to run everything in a single AWS account with IAM policies and tags doing the work of separation. I've moved away from that model every time, because an account boundary is a fundamentally stronger isolation primitive than a policy. A misconfigured IAM policy or a compromised credential in a single-account setup has a blast radius that spans your entire footprint: production, staging, and shared tooling all sit behind the same perimeter. Split those workloads across accounts and a mistake in one environment simply can't reach another. There's no policy to misconfigure your way around a hard account boundary. It also makes least privilege easier to reason about (permissions are scoped to what an account contains, not enumerated resource by resource), and it makes cost attribution trivial, since every dollar in an account belongs to a known team or workload without tagging discipline doing the heavy lifting.",[10,178,180],{"id":179},"a-pragmatic-account-structure","A pragmatic account structure",[15,182,183],{},"The structure I default to has a management account at the root purely for billing and organization-wide policy, a dedicated security and log-archive account that receives read-only copies of every other account's logs and holds no workloads of its own, a shared-services account for things like a container registry or CI runners that every environment needs but shouldn't own, and per-environment workload accounts (dev, staging, production) each isolated from the others. This isn't exotic. It mirrors what AWS Organizations and AWS Control Tower expect, and it keeps the mental model simple: if you can name the account, you know what's allowed to happen in it.",[10,185,187],{"id":186},"terraform-that-scales-with-you","Terraform that scales with you",[15,189,190,191,194],{},"None of this is manageable by hand once you have more than two or three accounts, so I express the entire structure in Terraform: reusable modules for the pieces that repeat across accounts (networking, IAM baselines, logging), and account- and environment-specific state, kept separate so a change in staging can never touch production's state file. Rollouts are automated end to end. There is no manual ",[93,192,193],{},"terraform apply"," from someone's laptop against a live account. If a change needs to happen, it happens through the same reviewed pipeline every time, which is what makes the account structure trustworthy in the first place. The isolation is only as good as the process that maintains it.",[10,196,198],{"id":197},"iam-and-logging-fundamentals","IAM and logging fundamentals",[15,200,201],{},"Inside that structure, I keep IAM roles scoped to least privilege by default (service roles get exactly the permissions their workload needs, nothing assumed \"just in case\"), and every account ships its CloudTrail and CloudWatch logs to the central log-archive account, where they're immutable from the perspective of the accounts that generated them. That combination means a compromised workload account can't erase its own evidence, and a review of \"what happened\" never depends on trusting the account under investigation.",[10,203,42],{"id":41},[44,205,206,209,212,215],{},[47,207,208],{},"Account boundaries isolate blast radius in a way no IAM policy alone can.",[47,210,211],{},"A small, purpose-built set of accounts (management, security, shared services, per-environment workloads) covers most real needs.",[47,213,214],{},"Terraform modules with separated per-account state keep the structure maintainable as it grows.",[47,216,217],{},"Centralized, immutable logging in a dedicated account is worth setting up before you need it.",{"title":60,"searchDepth":61,"depth":61,"links":219},[220,221,222,223,224],{"id":172,"depth":61,"text":173},{"id":179,"depth":61,"text":180},{"id":186,"depth":61,"text":187},{"id":197,"depth":61,"text":198},{"id":41,"depth":61,"text":42},"July 28, 2026",{},"\u002Fen\u002Fblog\u002Faws-multi-account-terraform",{"title":167,"description":60},"en\u002Fblog\u002Faws-multi-account-terraform","Account boundaries are your strongest isolation primitive. A practical structure for workload, security, and shared-services accounts, expressed entirely in Terraform.","LTfmcvJCsqfd8CfpI4ZRZlb88JZIaG7ei8utoEcV0p0",{"id":233,"title":234,"body":235,"category":156,"date":288,"description":60,"extension":70,"meta":289,"navigation":72,"order":290,"path":291,"seo":292,"stem":293,"summary":294,"__hash__":295},"blog\u002Fen\u002Fblog\u002Fkafka-topic-design-idempotency.md","Event-Driven Architecture with Kafka: Topic Design, Partitioning, and Idempotency",{"type":7,"value":236,"toc":281},[237,241,244,248,251,255,258,262,265,267],[10,238,240],{"id":239},"topics-are-contracts","Topics are contracts",[15,242,243],{},"I treat every Kafka topic as a contract between whoever produces to it and whoever consumes from it, not just a name in a config file. Naming conventions matter. A topic's name should tell you what domain event it carries and who owns it, not just what service happens to write to it today. Schema evolution needs to be an explicit, versioned decision rather than something that happens implicitly because a producer added a field. A schema registry with compatibility rules enforced at write time catches that before it becomes a consumer's runtime exception. Retention is the third decision I make deliberately for every topic: how long events need to live, and whether this topic is a transient message bus or a durable log that downstream systems might replay from the beginning. None of these are defaults I accept. They are decisions made once, up front, because changing them later means coordinating every producer and consumer at once.",[10,245,247],{"id":246},"partitioning-for-parallelism","Partitioning for parallelism",[15,249,250],{},"Partition count and partition key choice are where most of the design work lives, because they determine ordering guarantees and how evenly load spreads across consumers at once. Kafka only guarantees ordering within a partition, so events that must be processed in order relative to each other need the same partition key, typically an entity ID. Get the key wrong and you either lose ordering you needed or, more commonly, create a hot partition where one key's volume dwarfs the rest and a single consumer becomes the bottleneck for the whole topic. I size consumer groups to partition count deliberately, since a group can never have more active consumers than partitions, and extra consumers just sit idle.",[10,252,254],{"id":253},"idempotency-or-duplicates","Idempotency or duplicates",[15,256,257],{},"At-least-once delivery is the practical default for Kafka, which makes duplicates a certainty over a long enough time window, whether from producer retries or consumer rebalances. I design consumers to be idempotent rather than trying to guarantee exactly-once delivery at the broker level: every handler produces the same end state whether an event is processed once or twice, usually via an idempotency key checked against a store before a side effect runs. That same property is what makes replay safe. If I ever need to reprocess a topic from an earlier offset to recover from a bug, idempotent consumers make that a non-event instead of a data-corruption risk.",[10,259,261],{"id":260},"operating-for-throughput","Operating for throughput",[15,263,264],{},"Day to day, the two levers I watch are retention and consumer lag. Retention strategy (time-based, size-based, or compacted for topics that represent current state rather than a log of events) directly affects both storage cost and how far back a replay can reach. Consumer lag is the single most useful health signal a streaming system gives you: a consumer group falling behind its produced offsets is the earliest warning of a downstream bottleneck, well before anything looks broken from the outside.",[10,266,42],{"id":41},[44,268,269,272,275,278],{},[47,270,271],{},"Treat topic naming, schema evolution, and retention as explicit decisions, not defaults.",[47,273,274],{},"Partition key choice controls both ordering and load distribution. Get it wrong and you get hot partitions.",[47,276,277],{},"Design for at-least-once delivery: idempotent consumers make duplicates and replay safe.",[47,279,280],{},"Monitor consumer lag as your primary early-warning signal for the whole pipeline.",{"title":60,"searchDepth":61,"depth":61,"links":282},[283,284,285,286,287],{"id":239,"depth":61,"text":240},{"id":246,"depth":61,"text":247},{"id":253,"depth":61,"text":254},{"id":260,"depth":61,"text":261},{"id":41,"depth":61,"text":42},"July 14, 2026",{},3,"\u002Fen\u002Fblog\u002Fkafka-topic-design-idempotency",{"title":234,"description":60},"en\u002Fblog\u002Fkafka-topic-design-idempotency","The failure modes of streaming systems are quiet: duplicates, hot partitions, lagging consumers. Design decisions that prevent them, from partition keys to replay-safe consumers.","uSJT-1WsP13kwSv_fbSukGTtu-2dNB49dhZrzsftOBA",1790607906606]