Projects
Getting Claude Code onto shared compute
Repository · Getting Claude Code off my laptop and onto shared compute
When a CloudFormation deploy failed, someone usually messaged me to ask what went wrong. AWS isn’t everyone’s day to day, but troubleshooting that depends on one person doesn’t scale. I wanted developers to get a starting point without waiting for me.
Read more
What I built and why
I built cfn-investigator to run Claude Code headlessly on shared compute. The public repository is a simplified example, rebuilt from scratch and narrowed to CloudFormation. Give it a failing stack name and, optionally, a suspected commit. It reads stack state through the AWS MCP server and writes an analysis to CloudWatch Logs. There’s a place to add forwarding to Slack or another destination.
I chose CodeBuild because the job fit: clone source, run a script, and post the result somewhere. It gave me the shell tools and logging I needed without building a Lambda container or setting up Fargate networking. I used an Anthropic API key to avoid the Bedrock limitations described in the README. I was optimizing for a working prototype I could maintain alone.
Tradeoffs and lessons
There are compromises. AWS-managed ReadOnlyAccess is broader than this tool needs, and the tools are installed fresh on every run. The two-role split limits AWS calls through MCP; it doesn’t sandbox Claude, which can still access the build role’s credentials.
The useful result was giving someone enough context to take the next step. In the original implementation, the investigator identified a missing environment variable on a Fargate task definition. The developer added it and redeployed without needing to message me. The prompt also allows an “unsure” answer with ranked hypotheses. I’d rather get a useful shortlist than a confident wrong guess.
What I'd do next
If I built this again, I’d look at an agent SDK rather than running Claude Code headlessly. I’d want to spend less time on the plumbing around the agent. For a frequently used version of this prototype, I’d also narrow the IAM permissions and bake the tools into an image.
Running an open-weight model on AWS
I wanted to learn what it takes to run an open-weight model on AWS for my own coding workflow, including the infrastructure and costs of running it myself.
Read more
What I built and why
The starting point was one GPU instance running vLLM in Docker, deployed with CloudFormation. I connect from my laptop through an AWS Systems Manager Session Manager port-forwarding tunnel. The instance has no public IP or inbound security group rules, and inference stays inside my AWS account.
I chose vLLM partly because I wanted a container-based setup. One container on a GPU instance works for my own experiments. If a team depended on it, I’d look at ECS Managed Instances to manage it as a shared service. vLLM also provides the APIs my coding clients need and can serve concurrent requests against one copy of the model. For just experimenting on a single VM, running Ollama would have been easier.
Tradeoffs and lessons
I focused on getting the workflow working before choosing a final model. The default is gpt-oss-20b, and the project README notes noticeably weaker results than hosted Claude on multi-step tasks, especially with this default.
Even On-Demand GPU capacity was hard to find in the US regions I tried. I wanted to stay in North America for latency, but Canadian regions didn’t offer the instance types I needed at the time. I’d love to see enough capacity to make this reliable for everyday use.
And stopping the GPU doesn’t stop every charge: the network and storage still cost money while those resources exist.
What I'd do next
If I extended this into something a team depended on, I’d evaluate it against real coding tasks first.
Is it snowing in Hillsboro?
Repository · Serverless Weather Reporting with AWS Step Functions and CDK · I Rewrote My Step Function as a Durable Function
After moving from Portland to the suburbs, I noticed the Portland snow site wasn’t accurate for my area. I built my own version to answer one question: is it snowing in Hillsboro? I also wanted hands-on experience building a state machine with Step Functions.
Read more
What I built and why
The site gives visitors a simple YES or NO. Behind it, EventBridge Scheduler periodically starts a Step Functions workflow. The workflow compares the stored status with current conditions from OpenWeatherMap and updates the site when the answer changes. The current version stores status in Parameter Store and serves the static site through S3 and CloudFront. The infrastructure is defined in TypeScript with AWS CDK.
Tradeoffs and lessons
Step Functions gave me a visual workflow and managed retries and state tracking. I used direct service integrations where they fit, but HTML generation needed a Lambda function. My original attempt to write HTML directly through the S3 integration introduced extra quotation marks. Sending it as a buffer from Lambda worked.
The infrastructure has its own tradeoffs. I initially skipped CloudFront because the audience was local, then added it for HTTPS. Getting a certificate felt like a lot of infrastructure for such a small site. Step Functions also makes the execution easy to inspect, but passing JSON between states and defining error handling can make the CDK code verbose.
How it's evolved
I later rebuilt the workflow separately with Lambda durable functions to compare the developer experience. Writing the workflow in plain TypeScript felt more natural to me. I still preferred Step Functions’ visual debugging. Building the same project both ways gave me something concrete to compare.
I wouldn’t change much. The project has evolved as I’ve tried new AWS features, which was part of the learning.
More projects
- job-search-agent: A Strands/Bedrock AgentCore example for finding open roles. Read the write-up.
- More on GitHub