Site Reliability Engineer
One Inc · United States
Position Title: Site Reliability Engineer
Department: Technology
Reports To: Lead Site Reliability Engineer
FLSA Status: Salary, Exempt
Location: US, Remote
Overview: One Inc is hiring a Site Reliability Engineer to join our US SRE team. You will help keep our insurance payment platforms reliable, including ClaimsPay, Digital Payments, and the shared Azure infrastructure that supports them.
This is a hands-on role with a target balance of about 50% operations and incident response, and about 50% engineering and automation. You will work closely with product engineering, support, and cloud teams across US and international time zones.
Periodic on-call coverage is required.
Key Responsibilities:
- Own production reliability for multi-tenant Azure workloads, including AKS, App Services, VMs, networking, and data stores.
- Participate in a rotating on-call schedule to triage alerts, resolve incidents, communicate status clearly, and improve response through runbooks and automation.
- Build and maintain internal tooling such as Python scripts, Taskfiles, GitLab pipelines, and small services that support client onboarding, environment setup, migrations, and day-to-day operations.
- Improve observability with Prometheus/Grafana dashboards and alerts, log search, and runbooks tied to common failure modes.
- Work with development teams to diagnose and resolve reliability issues across application code, configuration, and infrastructure.
- Support operational programs such as ClaimsPay environment migrations, client cutovers, and production readiness reviews.
- Use AI-assisted engineering workflows (including Cursor) for investigation, documentation, and automation, while applying sound judgment for any production-impacting work.
Skills:
- Strong sense of responsibility and ownership.
- At least five years of experience in software engineering, SRE, DevOps, or system administration.
- Experience with process automation and scripting or software development languages. Python is preferred; Bash and PowerShell are also useful.
- Proficiency with Azure cloud services, including identity, networking, compute, storage, and monitoring.
- Experience administering Linux and Windows servers.
- Basic SQL and database skills, including reading schemas, writing SELECT queries, investigating data issues, and working with MySQL and/or SQL Server during production troubleshooting.
- Ability to diagnose and isolate issues across application, database, and infrastructure layers in development and production.
- Clear written communication for incident updates and operational documentation.
- Bachelor's degree in CS/Engineering or equivalent experience.
Preferred Qualifications
- Kubernetes / AKS
- CI/CD (GitLab CI, TeamCity, Octopus Deploy, or similar)
- Prometheus, Grafana, Elasticsearch/Kibana, or equivalent observability tools
- Deeper database operations experience, such as replication, backups, performance tuning, or schema migrations
- Infrastructure as Code experience with Terraform; Ansible experience is also useful
- Experience with AI coding assistants or similar tools (Cursor, Copilot, Claude, etc.)
- Experience in payments, fintech, or other regulated multi-tenant SaaS environments
Desired Traits:
Action Oriented, Growth Mindset, Positive Outlook, Problem Solver, Self-starter, Demonstrates Ethical Behavior, Strong Drive, Team Player, Supportive & Adaptable to Change, Exudes a commitment to Personal & Professional Development
How We Work:
- Target mix of about 50% operations and on-call support, and about 50% engineering and automation.
- We document work as we go in Confluence, GitLab, and Jira, and prioritize durable fixes over one-off firefighting.
- The team is adopting AI-assisted workflows for alert triage, RCA, and tooling. Prior AI experience is not required, but willingness to learn and contribute is.
Physical Demands:
The conditions herein are representative of those that must be met by an employee to successfully perform the essential functions of this job. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential job functions.
Environment:
Standard indoor office setting; exposure to computer screens.
Physical:
Requires repetitive motion. Substantial movements/motions of the wrists hands, and/or fingers. Sufficient mobility to work in an office setting; stand or sit for prolonged periods of time; operate office equipment including use of a computer keyboard, mouse, scanner and other tools as needed.
Vision:
See in the normal vision range with or without correction; vision sufficient to read computer screens and printed documents.
Hearing:
Ability to hear in the normal audio range with or without corrections.
Company Profile:
At One Inc, we empower insurers to meet policyholder expectations with choice, control, convenience, and continuity. Our mission is simple: to make every payment a promise kept.
The One Inc Insurance Payments Network seamlessly integrates multi-channel digital communications with inbound payment processing and outbound disbursement, delivering a frictionless experience for both premiums and claim payments. With over $120 billion in annual payments volume, we are proud to serve more than 300 insurance carriers, helping them honor their commitments instantly and securely.
Headquartered in Folsom, CA, One Inc offers competitive salaries, comprehensive benefits, including medical, dental, and vision insurance, a 401(k) plan, and a strong commitment to work-life balance. We believe in growing from within, promoting opportunities for career advancement across our team of 1,200+ dedicated "Onesters."
Join us in building the infrastructure that fulfills the promise of insurance.
One Inc is an equal opportunity employer and complies with all EEOC legislation in each jurisdiction it operates in.