Senior Site Reliability Engineer at Fingerprint

Company
Fingerprint
Employment type
Full-Time
Location
Worldwide
Posted
2026-09-29

About this role

Fingerprint empowers enterprises to detect and stop online fraud with the world’s most accurate device intelligence. We lead our industry with bleeding-edge identification capabilities and work on turning new ideas and discoveries in the fraud detection space into reality. Our customers range from innovative startups to leading enterprise companies, including Plaid, Dropbox, and Booking.com. Fingerprint is a globally dispersed, 100% remote company. We were named on on the 2026 Forbes Best Startup Employers list and ranked #803 on the 2026 Inc. 5000 list of America’s fastest-growing private companies. We have raised $77M and are backed by Craft Ventures ( Tesla, Facebook, Airbnb), Nexus Venture Partners ( Postman, Apollo.io, MinIO, Druva) and Uncorrelated Ventures ( Redis, Rollbar, Gradle). About the role Are you a systems-minded engineer who is happiest when production tells you something you didn't expect? Do you care less about how a system looks on a diagram than about how it behaves at 3am under load it wasn't designed for? Do you want to own reliability for a platform that answers millions of identification requests a day, where being wrong or being slow is a customer-visible event? If so, we have the perfect opportunity for you. We're looking for a Senior Site Reliability Engineer to join our Infrastructure team and take ownership of how our platform behaves in production. This is a hands-on engineering role, not an oversight one — you'll write code and infrastructure, own systems end to end, and be measured by whether the things you own stay fast, available, and predictable as we grow. You'll work across observability, incident response, capacity and performance, change safety, and the tooling that makes all of it routine. You'll define what "reliable" means for the critical paths you own, instrument them so we know before customers do, and partner with product engineering teams to make their services operable by design rather than by heroics. Responsibilities Own the reliability of core production systems end to end — you instrument them, set targets for them, operate them, and are accountable for how they behave under real traffic. Define and maintain SLIs and SLOs for the critical paths you own, wire them into dashboards and alerts, and use error budget burn as the evidence base for what gets fixed next. Drive alert quality: raise signal, kill noise, and close the gap where customers notice a problem before our monitoring does. Anomaly and correctness detection matter as much as uptime. Take a lead role in incident response — investigate systematically across service boundaries, restore service, and write postmortems that produce follow-ups people actually complete. Build secure, resilient, and cost-efficient infrastructure, with explicit attention to failure modes: timeouts and retries, backpressure and load shedding, graceful degradation, and blast radius containment. Do capacity and performance work with real data — load testing, profiling, saturation analysis, and headroom planning ahead of growth rather than after an incident. Improve change safety: progressive delivery, automated rollback, meaningful pre-production signal, and deployment practices that make shipping boring. Manage infrastructure through code and configuration (we primarily use Terraform), consistently applying patterns that align with our overall service architecture. Design, write, and ship software and developer-facing tooling that reduces toil and makes operating services straightforward for the engineers who own them. Run deliberate failure testing — game days and chaos exercises, staging first — to find the gaps and safe limits before customers do. Partner with product engineering teams on production readiness for new and high-risk services: capacity, failure modes, rollback plans, runbooks, and on-call handoff. Teach through review rather than gatekeeping. Participate in the on-call rotation, and improve it: better runbooks, clearer esca…

Apply for this Senior Site Reliability Engineer role

Keep browsing