Site Reliability · Cloud Platform · DevOps

I build the engineering organisations that make cloud platforms reliable, fast and affordable.

16+ years leading SRE, Cloud Platform and DevOps teams at Guidewire, Okta, Boeing and VMware — across India, the US, EMEA and APAC. I manage through managers, run platforms as products, and treat cost as a reliability metric.

Senior Manager, SRE (APAC) at Guidewire Bengaluru, India US B1 visa · valid to 2034

Impact

Results, not responsibilities

Every number below is tied to a specific role. Open the experience section to see where each one came from.

0→40
Org built from zero
Boeing Cloud Platform engineering
85%
MTTR reduction
120 → 18 min at Okta
$1M+
Cloud COGS savings
In 8 months at Boeing
99.99%
Availability sustained
Mission-critical SaaS at Okta

Experience

Filter by what you're hiring for

Pick a theme to highlight the outcomes that matter to your role. Select a company to expand it.

Earlier: VMware (2012–2018) — End User Computing Specialist, vCloud Air Engineer, IT Operations Analyst · Dell (2010–2012) — client technical support.

Approach

How I lead

Most reliability problems are really organisation problems. These are the habits I build into teams.

01

Platform as a product

Internal platforms get roadmaps, adoption metrics and a cost-to-serve number — the same discipline as customer-facing SaaS.

02

Leaders who build leaders

I manage through managers and build organisations that perform without me in the room.

03

AI-augmented, not AI-replaced

ML and LLMs where they remove toil and shorten incidents; humans stay accountable for judgment.

04

Cost is a reliability metric

An over-provisioned platform is a fragile business. FinOps belongs in engineering culture, not in a quarterly audit.

Try it

Error budget calculator

Every SLO is a business decision about how much unreliability you can afford. Here's what each one actually buys you.

Your SLO

Choose an availability target and a measurement window.

Availability target
Window
Allowed downtime
43m 12s
per 30 days at 99.9%

Budget burn

Enter the downtime you've had this window to see how much budget is left.

Past 75% burned, I'd freeze risky releases and spend the rest of the window on reliability work.

Toolkit

What I work with

Leadership & strategy

Managing managersGlobal orgsStrategic roadmappingFinOps / COGSVendor managementStakeholder management

SRE practice

SLIs / SLOsError budgetsBlameless postmortemsProduction readiness reviewsIncident commandCapacity planningToil reduction

Cloud & platform

AWSAzureGCPKubernetes (EKS/AKS/GKE)TerraformArgoCDSpinnakerGitOpsPolicy as Code

Observability

DatadogPrometheusGrafanaSplunkNew RelicDynatraceDistributed tracing

AI & automation

Predictive stabilityLLM-augmented triageSelf-healing runbooksML right-sizingPythonBash

Security & compliance

HIPAAGxPGDPRFDA 21 CFR Part 11RBACVulnerability management

Contact

Wrestling with reliability at scale?

I'm open to conversations about Director-level leadership in Site Reliability, Cloud Platform and DevOps engineering.

amith.rk05@gmail.com
Email me Connect on LinkedIn