AYN

Site Reliability Engineer (SRE)

Paragon · Tel Aviv-Yafo, Tel Aviv District · onsite

Paragon is on a mission to transform the world of cyber intelligence.

Based in Tel Aviv, our innovative team is made up of top-tier talent who are passionate about making an impact. At Paragon, you’ll find the freedom to think boldly, collaborate with purpose, and grow alongside a team united by a shared mission - striving for excellence, and always looking out for one another.

Responsibilities- Design and develop infrastructure and reliability solutions for complex customer environments, from architecture and development through production deployment.

- Develop automation for deployment, configuration, provisioning, and lifecycle management of our product and its infrastructure.

- Design and build observability solutions that provide deep visibility into system health, performance, reliability, and customer experience.

- Develop metrics, logs, events, dashboards, and alerting mechanisms to identify failures, understand system behavior, and proactively detect reliability issues.

- Build internal tools for debugging, diagnostics, troubleshooting, and operational visibility across application, container, system, and network layers.

- Investigate complex system and networking issues end-to-end, perform deep root cause analysis, and drive issues to resolution.

- Work closely with R&D, Deployment, and Support teams to improve observability, reliability, scalability, and operational efficiency.

Requirements- 3+ years of hands-on experience in SRE, DevOps, or infrastructure engineering, working with production environments.

- Strong software engineering skills in Python, including developing production-quality solutions and automation. Bash scripting experience.

- Strong Linux knowledge, including system troubleshooting, networking, processes, configuration, and kernel-level behavior.

- Strong understanding of networking and troubleshooting, including TCP/IP, DNS, routing, NAT, VPNs, firewalls, and proxies.

- Hands-on experience with Kubernetes, including cluster troubleshooting, networking, deployments, and Helm. Experience creating and maintaining Helm charts.

- Strong experience with observability, including metrics, logs, dashboards, alerting, and end-to-end troubleshooting. Experience with Prometheus, Grafana, Elastic, Fluentd/Vector, or similar technologies.

- Experience with infrastructure automation, using Ansible or similar technologies.

- Strong debugging, root cause analysis, and problem-solving skills, with the ability to troubleshoot complex issues across multiple infrastructure layers.

- Strong ownership, self-management, and technical curiosity, with the ability to independently investigate and drive complex technical problems to resolution.

- Strong communication and collaboration skills, working effectively across R&D, Deployment, Support, and other engineering teams.

Nice to have:

- Experience with Terraform / Terragrunt or similar Infrastructure-as-Code technologies.

Apply on the employer’s site