I build infrastructure platforms that scale, and the AI tooling that makes them effortless to use.
Senior Software Engineer at LinkedIn with 7+ years in infrastructure and distributed systems. I've spent the past 4 years building the control planes, resource managers, and tooling that 10,000+ engineers use to provision and operate 30+ infrastructure platforms. On that foundation I layer production AI as a force multiplier, to abstract complexity, simplify flows, and make efficiency gains through autonomous, regulated operations.
LinkedIn's greenfield Infrastructure-as-Code stack: the authoring, validation, and publishing layers plus a unified CLI, now used by 10,000+ engineers and agentic workflows to manage infrastructure. Shift-left validation catches misconfigurations at PR time, cutting apply-time failures ~90% across 5,000+ services (projected ~$5M/yr savings).
Led a 4-engineer team to build the first agentic layer that turns natural-language intent into validated, production-ready infrastructure, driven from the tools engineers already use (VS Code, Cursor, IntelliJ, Claude Code, GitHub Copilot CLI). It's pluggable by design: a new platform integrates by authoring one rules file instead of building its own agent (~17 days → ~1 day of dev).
A high-throughput (~30M QPS) service that tracks live utilization for a large distributed NoSQL document store. It became a core dependency of five workflows: database provisioning (cluster capacity checks), quota changes, deletion (live-traffic safety checks), usage-based quota recommendations, and infra-cost reporting.
Re-architected the manifest model into a backward-compatible, multi-file structure that lets a single manifest be split across repositories on demand, enabling file-level resource access control and bounding the context automation sees. Rolled it out fleet-wide to 154 services across 90+ teams with zero outages, executed by AI agents that raised the PRs, repaired CI, pinged owners, and tracked closure.
The resource-manager core at the heart of a new, centrally managed control plane, designed for pluggable platform onboarding. I brought the MySQL platform on as the first tenant, delivering self-serve provisioning with 70% fewer errors and 53% lower API latency; the pattern is now reused across 30+ platforms (Couchbase, Temporal, TiDB, and more).
The provisioning engine beneath the control plane, on Kubernetes and Crossplane, reconciling ~8M resources across 10+ platforms. I model complex, multi-step operations as resumable Kubernetes-object state machines; automating these previously manual flows cut provisioning time ~73%, and the engine is now extending into stateful cluster management.
AI woven through the platform's operations at several layers: an automated PR-review agent (on the GitHub Copilot SDK) that reviews every infrastructure change, a developer on-call agent that triages across design docs, wikis, Slack, tickets, and code, and user-facing assistant layers in the product UI and Slack. Backtested on real incidents, the on-call agent matched human judgment and saved 10–42 minutes per investigation.
A streaming framework that auto-rightsizes capacity for a large distributed document store, with pluggable recommendation engines, policy-driven enforcement, and near-real-time utilization tracking. Targets roughly $1.6M/yr of reclaimable capacity.
Extending the provisioning engine beyond logical resources into stateful cluster management. A ClickHouse proof-of-concept models the full lifecycle, bootstrap, scale, reshard, red/black failover, backup, and restore, as Kubernetes-object state machines, so teams can move off bespoke operators onto a shared platform.
Built three pillars of research infrastructure at Goldman Sachs: a database-agnostic graph database for fraud detection (Neptune, JanusGraph, Gremlin), a graph search service on Elasticsearch (autocomplete, fuzzy, weighted entity lookups), and a document/coverage store on Cassandra that cut response times ~35% over the legacy system.
Own the authoring, validation, and publishing layers of the IaC control plane and its unified CLI, plus the AI layer that helps engineers author, migrate, and operate infrastructure.
Built the resource-manager core of a new infrastructure control plane, now powering 30+ platforms, and a high-throughput (~30M QPS) resource-insights service that five provisioning and quota workflows depend on. Brought the MySQL platform on as the first tenant: 70% fewer errors, 53% lower API latency.
Built distributed graph, search, and document infrastructure for fraud detection and ML-powered research.
$ whoami
Amandeep Srivastava, Infrastructure & Platform Engineer
$ cat journey.log
4 years at LinkedIn building the control planes, resource managers, and tooling behind 30+ infrastructure platforms.
Before that, graph, search, and document infrastructure for fraud detection at Goldman Sachs.
Now layering AI on that foundation, so engineers, on-call, and platform partners work by intent instead of toil.
$ echo $EDUCATION
IIT (ISM) Dhanbad · B.Tech CS (Honors) · Class of 2019
$ echo $INTERESTS
distributed-systems, control-planes, infrastructure-as-code, agentic-AI, AI-for-infra
$ _