Site Reliability Engineering
منذ 4 أسابيع
New Cairo Cairo, 00, مصر
Konecta
دوام كامل
مجانًا عبر البريد الإلكتروني أو Google
احفظ هذه الوظيفة وحافظ على تنظيم بحثك
قم بإنشاء حساب مجاني لحفظ الوظائف وإنشاء التنبيهات والعودة إلى هذه القائمة من لوحة التحكم الخاصة بك.
مجانًا عبر البريد الإلكتروني أو Google
Mission:
Embed within the Kolibri team to learn the platform architecture and operating model,
then take ownership of run, support, and reliability (L1/L2) while escalating complex
issues (L3) to the platform team.
Core
Responsibilities:
1. Platform onboarding (first phase – critical) ● Join Kolibri squad(s) for several weeks ● Understand: ○ Control plane architecture ○ Agentic workflows & orchestration ○ Observability stack (logs, metrics, traces) ○ Deployment pipelines & environments ● Build operational knowledge of real use cases, not just infra 2. Run & Support (steady state) L1 / L2 ownership: ● Incident triage & resolution ● Monitoring platform health (SLA, latency, errors) ● Managing alerts & escalation flows ● Basic remediation (restart services, config fixes, rollback) Operational excellence: ● Improve runbooks ● Reduce MTTR ● Identify recurring issues (problem management) 3. L3 Interface with Kolibri team ● Escalate complex issues (design flaws, bugs, scaling limits) ● Provide structured feedback (logs, reproduction steps, impact) ● Act as bridge between Global IT and platform engineering Required
Skills:
SRE / Platform Ops fundamentals ● Incident management (ITIL mindset) ● Observability tools (Datadog, Prometheus, Grafana, etc.) ● Cloud environments (GCP/AWS/Azure) ● CI/CD understanding Strong plus (important for Kolibri): ● API-based systems & distributed architectures ● Event-driven systems / microservices ● Understanding of AI/LLM-based systems (at least operationally) ● Strong kubernetes knowledge, including cluster management an scaling ● Background in infrastructure as code ● Knowledge of GitOps
Responsibilities:
1. Platform onboarding (first phase – critical) ● Join Kolibri squad(s) for several weeks ● Understand: ○ Control plane architecture ○ Agentic workflows & orchestration ○ Observability stack (logs, metrics, traces) ○ Deployment pipelines & environments ● Build operational knowledge of real use cases, not just infra 2. Run & Support (steady state) L1 / L2 ownership: ● Incident triage & resolution ● Monitoring platform health (SLA, latency, errors) ● Managing alerts & escalation flows ● Basic remediation (restart services, config fixes, rollback) Operational excellence: ● Improve runbooks ● Reduce MTTR ● Identify recurring issues (problem management) 3. L3 Interface with Kolibri team ● Escalate complex issues (design flaws, bugs, scaling limits) ● Provide structured feedback (logs, reproduction steps, impact) ● Act as bridge between Global IT and platform engineering Required
Skills:
SRE / Platform Ops fundamentals ● Incident management (ITIL mindset) ● Observability tools (Datadog, Prometheus, Grafana, etc.) ● Cloud environments (GCP/AWS/Azure) ● CI/CD understanding Strong plus (important for Kolibri): ● API-based systems & distributed architectures ● Event-driven systems / microservices ● Understanding of AI/LLM-based systems (at least operationally) ● Strong kubernetes knowledge, including cluster management an scaling ● Background in infrastructure as code ● Knowledge of GitOps