Яндекс Метрика
Information Technologies about 3 hours назад · 0 откликов

Senior DevOps / Site Reliability Engineer - Healthcare / Voice / Observability

D
Deffic Компания
Information Technologies
Зарплата
По договорённости
На руки, до вычета НДФЛ · обсуждается
Локация
Не указано
Формат
Не указан
Добавлено
5 October 2026
О вакансии

Что предстоит делать

We are seeking a Middle+ / Senior DevOps / Site Reliability Engineer to join an international healthcare project focused on workflow and voice services for US health-system use. The platform supports medical voice services and workflow automation for healthcare systems in the US. The role focuses on production reliability: monitoring, incident response, SLOs, deployment safety, recovery, and uptime evidence. Candidates must have strong production experience in SRE / DevOps / Backend Operations, hands‑on experience with monitoring, logging, alerting, and tracing, real experience with Incident Response / Incident Management, and clear responsibility for uptime, SLA / SLO / service reliability. They should be able to define or work with SLOs and production dashboards, have a strong cloud infrastructure background, and experience with containers and deployment automation. Experience with Infrastructure as Code (IaC), strong CI/CD knowledge, Python or comparable scripting skills for operational troubleshooting and automation, diagnosing and resolving production failures, backup, recovery, and capacity testing, and the ability to communicate clearly during incidents and provide concise written updates is required. The role also involves defining service-level objectives, building and maintaining reliability dashboards, instrumenting metrics, logs, and traces, monitoring workflow, voice, and third‑party service dependencies, leading incident response during coverage hours, writing post‑incident reviews, improving deployment safety with rollback and staged releases, maintaining change records and release processes, testing capacity, backups, and disaster recovery, coordinating handover and escalation with engineering teams, providing accurate uptime and reliability evidence, and supporting continuous improvement of platform resilience. Strong advantage includes experience in regulated production environments, WebRTC / real‑time media / voice infrastructure, LLM‑backed services, handling third‑party provider outages, and familiarity with Grafana / Prometheus / Datadog or similar observability tools.

Требования

Что мы ждём от кандидата

  • Strong production experience in SRE / DevOps / Backend Operations
  • Hands‑on experience with monitoring, logging, alerting, and tracing
  • Real experience with Incident Response / Incident Management
  • Clear responsibility for uptime, SLA / SLO / service reliability
  • Experience defining or working with SLOs and production dashboards
  • Strong cloud infrastructure background
  • Experience with containers and deployment automation
  • Experience with Infrastructure as Code (IaC)
  • Strong CI/CD knowledge
  • Python or comparable scripting skills for operational troubleshooting and automation
  • Experience diagnosing and resolving production failures
  • Understanding of backup, recovery, and capacity testing
  • Ability to communicate clearly during incidents and provide concise written updates
Компания

О Deffic

D
Deffic Компания
Information Technologies
Готовы откликнуться?

Работодатель свяжется с вами через контакты из профиля.