Site Reliability Engineering

منذ 4 أسابيع

New Cairo Cairo, 00, مصر Konecta دوام كامل
Mission: Embed within the Kolibri team to learn the platform architecture and operating model, then take ownership of run, support, and reliability (L1/L2) while escalating complex issues (L3) to the platform team. Core

Responsibilities:
1. Platform onboarding (first phase – critical) ● Join Kolibri squad(s) for several weeks ● Understand: ○ Control plane architecture ○ Agentic workflows & orchestration ○ Observability stack (logs, metrics, traces) ○ Deployment pipelines & environments ● Build operational knowledge of real use cases, not just infra 2. Run & Support (steady state) L1 / L2 ownership: ● Incident triage & resolution ● Monitoring platform health (SLA, latency, errors) ● Managing alerts & escalation flows ● Basic remediation (restart services, config fixes, rollback) Operational excellence: ● Improve runbooks ● Reduce MTTR ● Identify recurring issues (problem management) 3. L3 Interface with Kolibri team ● Escalate complex issues (design flaws, bugs, scaling limits) ● Provide structured feedback (logs, reproduction steps, impact) ● Act as bridge between Global IT and platform engineering Required

Skills:
SRE / Platform Ops fundamentals ● Incident management (ITIL mindset) ● Observability tools (Datadog, Prometheus, Grafana, etc.) ● Cloud environments (GCP/AWS/Azure) ● CI/CD understanding Strong plus (important for Kolibri): ● API-based systems & distributed architectures ● Event-driven systems / microservices ● Understanding of AI/LLM-based systems (at least operationally) ● Strong kubernetes knowledge, including cluster management an scaling ● Background in infrastructure as code ● Knowledge of GitOps