Lead Software Engineer · AI infrastructure · Ahmedabad, India

Agents don't need to be smarter. They need something true to stand on.

Ten years of backend and platform work. Since 2025 I have built the deterministic layer that lets AI agents do real engineering inside a real company: a compiled graph of the whole system, a governed delivery pipeline, and the evals and telemetry that keep both honest. Below: 6 tools you can install right now, and three essays on how they work.

real commands · real output captured 2026-09-04
every line is the command's real output on my machine, nothing typed by hand · click a tab to run another
10
years shipping
6
tools in public
3
packages on npm
13
coding harnesses served
3 Sep 2026
last public push

The work

Ten years, four bets.

Every job on this page was a bet that the boring, structural work would matter more than the flashy work. It kept paying. Scroll: the stage on the right follows the story.

Act 01 · 2016 → 2021

A founding-team seat, through the 0→1.

Joined a logistics-AI startup as an intern and stayed through its 0→1, from the first customers to a platform other teams built on. Event-driven microservices on GCP and Kubernetes, built while the product was still finding its shape.

What it taught me: a system you can reason about beats a system you can only observe.
GCPKubernetesevent-drivenRubyNode.jsElasticsearch
Act 02 · 2021 → 2024

Gave a building an API.

Senior engineer on an access-control and smart-buildings platform, where the software ends at a door. I drove R&D on the in-house IoT gateway (AWS IoT Core, MQTT, over-the-air updates), integrated elevator destination dispatch and third-party readers, and built the I/O rule engine: if-this-then-that automation for doors, readers and alarms. Around it I shipped the public APIs and webhook ecosystem partners build on, rate-limited with Redis, the video integrations, and a global dashboard for security operations centres at large multinationals.

What it taught me: when the failure mode is a locked door, "eventually consistent" is not a feature.
AWS IoT CoreMQTTOTAelevator dispatchrule enginepublic APIswebhooksRedis
Act 03 · 2022 → 2024

Moved a live product off Heroku, onto AWS.

Led the migration of 12+ live services from Heroku to AWS Fargate, with Terraform owning the infrastructure and GitHub Actions owning every deploy.

What it taught me: the migration nobody notices is the one you planned twice.
AWSECS FargateTerraformGitHub ActionsCI/CDPostgreSQL
Act 04 · 2025 → now

Gave agents something true to stand on.

Lead engineer for an AI platform across a fleet of more than a hundred services. Every AI coding tool fails the same way inside a real company: not at writing code, but at knowing the organisation. Which service consumes which event. What breaks three repos away when a column is renamed. Which rules are non-negotiable.

So I stopped asking models to guess and compiled it: a graph of the whole system that answers structural questions in milliseconds and re-checks itself against live source, a spec-driven pipeline agents operate inside rather than around, and evals and telemetry that keep both honest. The engineering team uses it every day.

What it taught me: determinism before generation. Look facts up; spend the model on judgment.
MCPClaude Codeknowledge graphsevalsPythonTypeScript

Tools I ship

Tools pulled out of daily agent work. Install any of them right now.

5 of them MIT, one free for personal use. Zero dependencies wherever that was possible. Each one exists because I needed it on a Tuesday.

skills: one install, every harness. One clone fanning out to Claude Code, Codex, Cursor, OpenCode and more

skills

npm · MIT · v0.1.2 · 10 skills · 18 agents · 13 harnesses

Skills and agents for AI coding assistants, installed into Claude Code, Codex, Cursor, OpenCode, Copilot, Gemini CLI and more from one central clone. Symlinked, so update is one git pull and every harness sees it. Agents transpile to each harness's native format. Includes visual-verify: the agent renders what it built and looks at it before claiming it works.

$ npx @vimoxshah/skills
TokenFlow dashboard: tokens per day across Claude Code, Codex and Cursor

tokenflow

npm · Homebrew · MIT · v1.1.2 · 0 dependencies

See where your AI tokens actually go. Reads the logs your tools already write and turns them into honest analytics across Claude Code, Codex, OpenCode, Cursor, Cline and Hermes: 10 adapters, 145 tests, CI on three operating systems. Dashboard, native menu-bar app and CLI. An unpriced model shows up as null with a fix, never as a silent $0. Nothing leaves your machine.

$ npx @vimoxshah/tokenflow demo
Clockwork: a month calendar with scheduled agent jobs, each showing its agent profile, repo and budget

clockwork

macOS app · Homebrew · v0.4.0 · free for personal use

The calendar where your AI agents show up for work. Book a recurring job for Claude Code, Codex, OpenCode or Hermes; it runs unattended inside an OS-sandboxed git worktree that never touches main and cannot read your SSH keys, then files a report you can actually read: what it did, what it skipped and why, what it cost. 13 agent profiles ship with it. Hard caps on dollars, turns and wall-clock.

$ brew tap vimoxshah/clockwork && brew install --cask clockwork
Animated demo: scrubbing through a recorded agent session in the replay player

claude-session-replay

npm · MIT · v1.0.0 · one HTML file out

A video player for what your agent actually did. Turns a session transcript into one self-contained HTML file you can scrub, step through and hand to a teammate. No server, no dependencies, and nothing leaves your machine until you share the file.

$ npx claude-session-replay
Diagram: tasks routed across model-pinned subagent lanes by task class

claude-router

skill + 5 lane agents · MIT

Route each task to the cheapest Claude tier that can do it well, through model-pinned subagents. Volume to Haiku, building to Sonnet, review and hard problems to Opus, judgment to the top tier. Only write lanes touch code, and every diff is verified by a different model than the one that wrote it. A task that fails twice escalates with its failure log, never a blind third retry.

$ npx @vimoxshah/skills --bundles routing
Diagram: the plan, dispatch, execute, return and verify loop between Claude and Codex

claude-codex-orchestrator

skill + Codex profiles · MIT

An orchestrator and an executor split across two vendors. Claude plans, decomposes and judges; Codex writes the code once the spec is frozen and makes no design calls of its own. Bounded work packets, three executor tiers, an escalation path, and the accepting verdict never comes from the model that wrote the diff.

$ npx @vimoxshah/skills --bundles routing

How I think

Six things I believe, each with the place it shows up.

01

Determinism before generation

Wherever a model could be asked, a lookup is tried first. Models are for judgment. Retrieval, blast radius and contracts are facts, and facts should be looked up, not generated.

Where it shows: claude-router is a table, not a vibe: task class → model tier, and a task escalates only on evidence, after it has failed twice.
02

The graph reports its own staleness

The first objection to any knowledge layer is "it will rot." So the layer checks itself: structural answers are re-verified against live source before anyone acts on them.

Where it shows: this page. Every number on it is generated from GitHub, npm or a dated résumé block, and the stamp in the dock opens the record of which. The essay on the graph that checks itself has the mechanism.
03

Attention is a budget

Published research puts the ceiling for reliable agent tool selection at roughly thirty tools. Fewer, higher-leverage capabilities beat a long menu, and every new one has to justify its slot.

Where it shows: skills ships ten skills, not a hundred, and each one does one job. The essay on the attention ceiling covers the consolidation.
04

Review that argues back

A reviewer that agrees with everything is decoration. Review is structured as disagreement: independent read-only lanes, then a lane whose only job is to refute a passing verdict.

Where it shows: in claude-codex-orchestrator the accepting verdict never comes from the model that wrote the diff. Same-model review re-runs the blind spots that produced the code.
05

Honesty is a feature

Impact claims stay unpublished until they are measured. Telemetry ships structure, never content. "I could not look" is never reported as "nothing found."

Where it shows: tokenflow shows an unpriced model as null with a configure action, never a silent $0, and keeps measured cost apart from estimates. clockwork's reports say what the agent skipped, and why.
06

Not building is a decision too

The most useful architecture record is often the thing that was evaluated and declined, with the reasoning written down where the next person will find it.

Where it shows: claude-session-replay has no server, no account and no dependencies, on purpose. The essay on the agent that cannot merge is about a capability removed rather than forbidden.

Now · September 2026

Where I am, and how I got here.

Portrait of Vimox Shah

Right now I am shipping the tools above and working on the unglamorous part of autonomous agents: making their reports honest, so that "I could not look" never comes back as "nothing found." The next thing I want to publish is a head-to-head benchmark of agents with and without a compiled graph of the codebase, with real error bars.

I have spent ten years on backend and platform work: a founding-team seat at a logistics startup through its 0→1, then access control and smart buildings, where I gave the hardware one gateway and one public API and led the move from Heroku to AWS, and since 2025 have led the AI platform work described above.

  • 2025 → nowLead Software Engineer, Genea. AI platform and hardware systems. Act 04 above.
  • 2022 → 2024Senior Software Engineer II, Genea. Heroku → AWS onto Fargate with Terraform and GitHub Actions; in-house IoT gateway R&D (AWS IoT Core, MQTT, Mender OTA); KONE elevator destination dispatch; the public APIs and webhook ecosystem with Redis rate limiting; Webex by Cisco and AVA Cloud VMS video integrations.
  • 2021Senior Software Engineer, Genea. Global dashboard for security operations teams at large multinationals; the I/O rule engine, time-based access groups and key control; Wavelynx readers and Schindler elevator destination dispatch.
  • 2016 → 2021Backend Software Engineer, Shipmnts. Founding team, intern to engineer, through the 0→1. Architected the AI product that cut users' day-to-day operational effort by ~70%; event-driven microservices on GCP and Kubernetes; live container tracking, ERP integrations, G Suite add-ons; mentored interns.
  • EducationB.E. Computer Engineering, Dharmsinh Desai Institute of Technology.
  • CredentialsGoogle Cloud Architect track, 4 certifications · Redis University RU101.
  • LanguagesEnglish and Hindi, fluent · Gujarati, native.
AI platform
Model Context ProtocolClaude Codeagent architectureknowledge graphsontology engineeringeval harnessesspec-driven developmentRAG
Backend
PythonTypeScriptNode.jsRubySQLDjangoFlaskRailsExpressRESTwebhooksrate limitingCelery · Sidekiq
Data & cloud
PostgreSQLMongoDBRedisElasticsearchSQLiteAWS · ECS FargateLambdaSNS · SQSS3CloudWatchGCPDockerKubernetesTerraformGitHub Actions
IoT & architecture
MQTTAWS IoT CoreMender OTAevent-drivenmicroservicesmulti-tenant SaaS

Hire me

If you are building the place where agents do real engineering work, I would like to talk.

Agent infrastructure, evals and deterministic retrieval; cloud platforms and the migrations that keep them alive; hardware that has to answer to software. Or a team that needs someone who ships the whole loop and measures it honestly.

vmoksh.shah179@gmail.com · Ahmedabad, India (UTC+5:30)