Razorpay · Bengaluru

Site Reliability Engineer

Posted 24 Sep 2026 · found on Razorpay's careers page 24 Sep 2026, 15:16 UTC

SRE role building scalable payment infrastructure at Razorpay, focusing on reliability, automation, and AI-driven observability.

Stack: go, python, java

Apply on Razorpay's site Ask for a referral

We found this role on Razorpay's own careers page. Members get roles like it as soon as we find them, and paid plans email the ones that match their resume.

Get new roles first — free

About the role

<div class="content-intro"><p>Razorpay is one of India’s leading full-stack financial technology companies, powering the way businesses move, manage, and grow money. Founded in 2014 by Harshil Mathur and Shashank Kumar with a simple vision - to simplify payments for Indian businesses - we’ve since grown into a fintech powerhouse driving India’s digital payment revolution.</p> <p>Razorpay powers millions of businesses with a smarter, scalable stack that goes beyond transactions to help them truly build and grow.</p> <p>From building AI-native agentic payments, to AI-assisted fraud detection and real-time risk intelligence to automated reconciliation, smart payouts, and predictive financial insights, we are embedding intelligence across our stack to make money movement faster, safer, and more efficient. In close collaboration with ecosystem partners - including banks, networks, regulators - we are pioneering industry-first solutions that are shaping the next era of fintech</p> <p>Across India, Singapore and Malaysia, our products span everything from seamless checkouts to payroll automation - powering a fintech ecosystem that’s redefining how money moves across Asia.</p> <p>Today, that ecosystem supports everyone from early-stage startups to some of India’s largest enterprises, enabling them to accept, process, and disburse payments at scale while expanding into new ways of managing money more efficiently.</p> <p>Our scale speaks volumes: Razorpay processes $180+ billion in annualized transactions, powering leading businesses like Airbnb, Facebook, WhatsApp, Airtel, CRED, BookmyShow, Zomato, Swiggy, Lenskart, Mirae Asset Capital markets, Indian Oil, National Pension Scheme - and over 100 of India’s unicorns. With strong roots in India and growing operations in Southeast Asia, we are shaping the next chapter of financial technology across the region.</p> <p>We are backed by global investors including GIC, Peak XV Partners (formerly Sequoia Capital India &amp; SEA), Tiger Global, Ribbit Capital, Matrix Partners, MasterCard, and Salesforce Ventures, having raised over $740 million to date. Strategic acquisitions - including Ezetap (POS and offline payments), Curlec (Malaysia expansion), BillMe (digital invoicing), and POP (rewards-first UPI) - along with earlier moves in fraud prevention, payroll, and lending, have further strengthened our platform and widened our footprint across Asia.</p> <p>But what truly sets Razorpay apart is our culture. At Razorpay, ownership is our oxygen - you own what you build, with no micromanagement or red tape, just the runway to make your ideas fly. Learning is a lifestyle - if you’re curious, you’ll feel at home here. People &gt; Pedigree - we hire for attitude, hustle, and hunger more than degrees. Transparency thrives over titles - this is where interns question CXOs and CXOs say “thank you.” Guided by our values of Customer First, Autonomy &amp; Ownership, Agility with Integrity, Transparency, Challenging the status quo and a strong belief that Razorpay grows with Razors,&nbsp; you’ll be part of a 3000+ strong team building not just products, but the financial infrastructure of the future.</p></div><p><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">About the Role</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Site Reliability Engineering (SRE) at Razorpay combines software and systems engineering to build&nbsp;and run large-scale, massively distributed, fault-tolerant systems that power India’s digital payment&nbsp;infrastructure. SRE ensures that Razorpay’s services — both our internally critical and our externally-&nbsp;visible systems — have the reliability, uptime, and performance appropriate to the needs of millions of&nbsp;businesses and their customers. Given that every transaction on our platform represents real money&nbsp;movement, the stakes for reliability are exceptionally high.&nbsp;As an SRE, you will keep an ever-watchful eye on system capacity, performance, and latency across a&nbsp;stack processing $180+ billion in annualized transactions. Much of our software development focuses&nbsp;on optimizing existing systems, building infrastructure, and eliminating manual work through&nbsp;automation. On the SRE team, you’ll have the opportunity to manage the complex challenges of scale&nbsp;that are unique to a high-throughput fintech platform, while using your expertise in coding, algorithms,&nbsp;complexity analysis, and large-scale system design.&nbsp;You will increasingly leverage AI and machine learning to transform how we operate — from AI-&nbsp;assisted incident detection and response, to ML driven anomaly detection, predictive capacity planning,&nbsp;and intelligent alerting that reduces noise and surfaces real issues before they impact customers.&nbsp;SRE’s culture of intellectual curiosity, problem solving, and openness is key to its success. Our&nbsp;organization brings together people with a wide variety of backgrounds, experiences, and perspectives.&nbsp;We encourage them to collaborate, think big, and take risks in a blame free environment.</span></p> <p><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">The SRE team at Razorpay supports services that are foundational to running production at scale — including payment routing, rate limiting, shard management, traffic shaping, and the reliability of core&nbsp;payment, payout, and reconciliation systems. These platforms implement safety guardrails, risk&nbsp;controls, regulatory compliance mechanisms, accounting integrity, and dynamic load balancing for&nbsp;stateful services. Behind everything our customers see is the architecture built by the Infrastructure&nbsp;team to keep it running, and we are proud to be our engineers’ engineers.</span></p> <p><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Responsibilities</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Develop strong, influential relationships with multiple stakeholders across the Site Reliability&nbsp;Engineering, Developer, and Product organizations.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Serve as an expert on particular fields of knowledge related to rate limiting, sharding, traffic&nbsp;management, and large-scale distributed system reliability for Razorpay’s payment platform.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Develop plans and lead projects that evolve our production systems and their reliability —&nbsp;spanning capacity planning, failure-mode analysis, and architectural improvements for high-&nbsp;throughput, low-latency payment infrastructure.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Design, build, and operate automation tooling and self-healing infrastructure that eliminates&nbsp;manual toil and improves system resilience across the payments stack.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Leverage AI and ML to build intelligent observability — anomaly detection, predictive alerting,&nbsp;noise reduction, and correlation engines that surface real issues before they impact customers.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Drive AI-assisted incident management: use LLM-based tooling for root-cause analysis, incident&nbsp;summarization, runbook generation, and post-incident learning to reduce mean-time-to-&nbsp;resolution (MTTR).</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Build and maintain reliability for AI-native and AI-assisted products — including agentic&nbsp;payments, AI-driven fraud detection, and real-time risk intelligence — ensuring the infrastructure&nbsp;supporting ML models and inference pipelines is production-grade.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Identify internal and external opportunities to improve systems, including evaluating emerging&nbsp;AI/Ops (AIOps) tools, observability platforms, and reliability patterns.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Provide on-call support for onboarded services, participating in incident response, blameless&nbsp;postmortems, and driving follow-up remediation.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Mentor engineers across the organization on SRE practices, reliability thinking, and the effective&nbsp;use of AI tooling in day-to-day operations.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Champion SRE best practices — service-level objectives (SLOs), error budgets, toil reduction,&nbsp;and data-driven reliability decisions — across engineering teams.</span></p> <p><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Minimum Qualifications</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Bachelor’s degree in Computer Science, a related technical field, or equivalent practical&nbsp;experience.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• 5 years of experience in software engineering, systems engineering, or site reliability engineering.&nbsp;3 years of experience with site reliability engineering focused on building and maintaining&nbsp;scalable, reliable systems.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• 3 years of experience in software design and architecture, including distributed systems and&nbsp;backend services.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Proficiency in at least one programming language (Go, Python, Java, or similar) with the ability&nbsp;to write production-quality code and automation.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Working knowledge of AI-assisted development and operations tooling — including experience&nbsp;using LLM-based copilots for code generation, debugging, and incident response workflows.</span></p> <p><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Preferred Qualifications</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Master’s degree in Computer Science or a related technical field.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• 5+ years of experience in large-scale distributed systems, preferably in a high-throughput,&nbsp;mission-critical domain (payments, fintech, e-commerce, or cloud infrastructure).</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Experience in a Site Reliability Engineering role at scale, with demonstrated impact on system&nbsp;availability, latency, or operational efficiency.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Experience leading complex process improvement projects and influencing and managing&nbsp;stakeholder relationships across engineering and product.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Experience in troubleshooting and debugging complex distributed systems under production&nbsp;pressure.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Hands-on experience with AIOps, ML-based anomaly detection, or building intelligent alerting /&nbsp;observability platforms.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Experience with LLM-powered tooling for operations — such as AI-driven incident response,&nbsp;automated root-cause analysis, or intelligent runbook automation.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Familiarity with building and operating infrastructure for ML/AI workloads — including model&nbsp;serving, inference pipelines, and the reliability of AI-native products.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Knowledge of chaos engineering, fault injection, and resilience testing for distributed systems.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Experience with cloud-native infrastructure (Kubernetes, service mesh, cloud platforms) and&nbsp;infrastructure-as-code.</span></p> <p><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">AI Competency</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">At Razorpay, AI is not a bolt-on — it is embedded across our stack and across how we operate. We are&nbsp;building AI-native agentic payments, AI-assisted fraud detection, and real-time risk intelligence. We&nbsp;expect SREs to be fluent in leveraging AI to operate more intelligently, and to build the reliability&nbsp;foundations for AI-powered products. The following AI competencies are expected for this role:</span></p> <p><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">AI for Operations (AIOps)</span><br><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Experience using LLM-based tools (e.g., coding copilots, AI assistants) in day-to-day SRE work&nbsp;— for code generation, debugging, runbook generation, and incident triage.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Ability to build or integrate ML-driven anomaly detection and intelligent alerting that reduces&nbsp;alert noise and surfaces genuine issues before customer impact.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Experience with AI-assisted incident response — using LLMs for log correlation, root-cause&nbsp;hypothesis generation, incident summarization, and automated postmortem drafts.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Familiarity with predictive capacity planning and forecasting using ML models to anticipate&nbsp;resource needs and prevent bottlenecks.</span></p> <p><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Reliability for AI Systems</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Experience building and operating reliable infrastructure for ML/AI workloads — including model&nbsp;serving, inference pipelines, feature stores, and the data plumbing that feeds them.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Understanding of the unique failure modes of AI systems — model drift, data quality&nbsp;degradation, inference latency spikes, and hallucination / output reliability — and how to build&nbsp;guardrails and monitoring for them.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Ability to define SLOs and error budgets for AI-native products, including agentic payment flows&nbsp;and AI-driven risk decisioning, where the definition of “correctness” is more nuanced than</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">traditional services.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">AI-Native Mindset</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Strong opinion on where AI helps versus where it adds risk in production operations — knowing&nbsp;when to trust AI-driven decisions and when to keep a human in the loop.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Eagerness to experiment with emerging AI tooling (agentic frameworks, LLM-based ops&nbsp;assistants, autonomous remediation) and evaluate their applicability to Razorpay’s reliability&nbsp;challenges.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Ability to mentor teammates on effective and responsible use of AI in engineering workflows,&nbsp;including prompt hygiene, output verification, and avoiding over-reliance on automated&nbsp;decisions in safety-critical paths.</span></p> <p><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">What You’ll Work On</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Reliability of Razorpay’s core payment processing, payout, and reconciliation systems — the&nbsp;infrastructure powering $180+ billion in annualized transactions.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• AI-native and AI-assisted products — including agentic payments, AI-driven fraud detection, and&nbsp;real-time risk intelligence platforms — ensuring they are production-grade and trustworthy.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• High-throughput, low-latency infrastructure: rate limiting, sharding, traffic management, and&nbsp;dynamic load balancing for stateful services.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Intelligent observability: building the next generation of monitoring, alerting, and incident&nbsp;management with AI woven in to reduce noise and improve signal.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Automation and toil elimination: building self-healing systems that detect, diagnose, and&nbsp;remediate issues without human intervention where safe.</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">• Capacity and performance: ensuring the platform can scale to support Razorpay’s growing&nbsp;transaction volume and expanding footprint across India and Southeast Asia.&nbsp;</span></p> <p><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Equal Opportunity</span><br><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Razorpay is an equal opportunity employer. We celebrate diversity and are committed to building an&nbsp;inclusive environment for all Razors. We hire for attitude, hustle, and hunger over mere degrees, and&nbsp;we believe the best teams are made of people with different backgrounds, experiences, and&nbsp;perspectives. If you’re excited about building the financial infrastructure of the future, we’d love to hear&nbsp;from you</span></p><div class="content-conclusion"><div class="gmail_default"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span id="m_2989597180337834284gmail-m_4972969247898306296gmail-docs-internal-guid-3a65a3c2-7fff-88ff-9e94-8ab11a050d04">Razorpay believes in and follows an equal employment opportunity policy that doesn't discriminate on gender, religion, sexual orientation, colour, nationality, age, etc. We welcome interests and applications from all groups and communities across the globe. </span><br></span></div> <div class="gmail_default"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">&nbsp;</span></div> <div class="gmail_default"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span id="m_2989597180337834284gmail-m_4972969247898306296gmail-docs-internal-guid-3d0a9248-7fff-a2fd-6fa3-5026f85768d7">Follow us on <a href="https://www.linkedin.com/company/razorpay/mycompany/" target="_blank" data-saferedirecturl="https://www.google.com/url?q=https://www.linkedin.com/company/razorpay/mycompany/&amp;source=gmail&amp;ust=1660290870959000&amp;usg=AOvVaw0f6sCrv8Ce3IHBvN2Sev8Z">LinkedIn</a> &amp; <a href="https://twitter.com/Razorpay" target="_blank" data-saferedirecturl="https://www.google.com/url?q=https://twitter.com/Razorpay&amp;source=gmail&amp;ust=1660290870959000&amp;usg=AOvVaw0ViuP9uutFg1qFCm2nHeh1">Twitter</a></span></span></div></div>

More at Razorpay