Glance · Bangalore

SDE IV - GPU Engineer

Posted 9 Oct 2025 · found on Glance's careers page 28 Sep 2026, 15:21 UTC

Lead GPU inference optimization for diffusion and transformer models at scale.

Stack: cuda, triton, c++, nccl, nvlink, pcie, tvm, xla, mlir, rocm, tpu

Apply on Glance's site Ask for a referral

We found this role on Glance's own careers page. Members get roles like it as soon as we find them, and paid plans email the ones that match their resume.

Get new roles first — free

About the role

<div class="content-intro"><p><strong>Glance</strong></p> <p>Glance is an intelligent shopping agent, redefining the commerce experience. Powered by proprietary agentic intelligence and generative AI, Glance delivers a hyper-personalized consumer experience across mobile and TV — shaping the new era of shopping. Glance is operated by Glance InMobi Pte. Ltd., a non-consolidated subsidiary of global technology leader InMobi, and is backed by Mithril Capital, Google and Jio Platforms. To learn more, visit glance.com.</p> <p><strong>InMobi&nbsp;</strong></p> <p>InMobi Group is a global technology company shaping the future of agentic commerce and advertising. Through its ecosystem of businesses — including InMobi Advertising and flagship consumer platform Glance — InMobi leverages data, machine learning, and generative AI to help brands reach audiences more precisely and consumers discover products more intuitively. Glance, which is pioneering new models of agentic commerce, is owned and operated by Glance InMobi Pte. Ltd., a non-consolidated subsidiary of InMobi Pte. Ltd.</p> <p><strong>InMobi Advertising</strong></p> <p>InMobi Advertising, part of global technology company InMobi, is an agentic advertising platform helping brands and merchants achieve their business outcomes. Through its proprietary intelligence, AI-led solutions, and vast consumer reach — including flagship consumer platform Glance — InMobi Advertising delivers the omnichannel performance defining what's next in advertising and commerce. Glance is owned and operated by Glance InMobi Pte. Ltd., a non-consolidated subsidiary of InMobi Pte. Ltd. To learn more, visit advertising.inmobi.com.</p></div><p><strong>About the Role</strong></p> <p>As a <strong>GPU Systems Engineer</strong>, you’ll lead design and optimization efforts across our GPU inference stack.<br>You will architect the libraries and runtime systems that enable <strong>Stable Diffusion, multimodal transformers</strong>, and emerging <strong>video generation</strong> models to run efficiently at scale.</p> <p>You’ll guide cross-functional teams, influence hardware selection, and set the technical vision for GPU optimization practices across the company.</p> <p><strong>Key Responsibilities</strong></p> <ul> <li>Architect <strong>high-performance inference runtimes</strong>, kernel dispatchers, and memory planners for large diffusion and transformer workloads.</li> <li>Lead investigations into <strong>cross-GPU performance bottlenecks</strong>, communication overheads, and scheduling inefficiencies.</li> <li>Drive <strong>multi-GPU parallelism strategies</strong> — model, pipeline, and tensor parallelization.</li> <li>Establish company-wide <strong>GPU optimization standards, tooling, and SLIs</strong>.</li> <li>Collaborate with research to design scalable implementations of novel architectures.</li> <li>Mentor engineers in profiling, tuning, and low-level optimization.</li> <li>Partner with hardware vendors and infra teams to maximize cluster utilization.</li> </ul> <p><strong>Required Qualifications</strong></p> <ul> <li>5+ years in high-performance computing, GPU runtime systems, or ML infrastructure.</li> <li>Proven expertise in <strong>CUDA / Triton / C++</strong>, with deep understanding of GPU scheduling, occupancy, register usage, and tensor cores.</li> <li>Experience building and maintaining <strong>distributed inference</strong> or training systems.</li> <li>Ability to <strong>design abstractions</strong> balancing flexibility and performance.</li> <li>Strong knowledge of <strong>NCCL</strong>, NVLink, PCIe, and interconnects.</li> <li>Familiar with <strong>profiling automation</strong> and performance dashboards.</li> <li>Excellent technical leadership and mentoring capabilities.</li> </ul> <p><strong>Preferred Qualifications</strong></p> <ul> <li>Background in <strong>compiler-aided optimization</strong> (TVM, XLA, MLIR, Triton).</li> <li>Experience tuning <strong>Stable Diffusion or transformer</strong> inference pipelines.</li> <li>Exposure to <strong>heterogeneous compute backends</strong> (AMD ROCm, TPU, ASICs).</li> <li>Experience working with <strong>hardware–software co-design</strong> initiatives.</li> <li>Open-source or research contributions in GPU optimization</li> </ul> <p>&nbsp;</p><div class="content-conclusion"><p><span class="ui-provider ee bcu auz bcv bcw bcx bcy bcz bda bdb bdc bdd bde bdf bdg bdh bdi bdj bdk bdl bdm bdn bdo bdp bdq bdr bds bdt bdu bdv bdw bdx bdy bdz bea">"<em>Glance collects and processes personal data such as your name, contact details, resume and other information that may contain personal data for the purpose of processing your application. Glance utilizes Greenhouse, a third-party platform. Please review Greenhouse's Privacy Policy to understand how the data collected from you is processed and managed. By clicking on 'Submit Application', you acknowledge and agree to the above privacy terms. Should you have any privacy concerns, you may contact us through the details mentioned in your application confirmation email."</em></span></p></div>

More at Glance