Hi, I'm Yazan.
I'm an AIOps Engineer. I specialize in fusing AI Agent Engineering with MLOps to transform infrastructure, turning manual operations into intelligent, self-healing systems powered by SRE and DevOps best practices.
Over the past 6+ years, I've architected and scaled cloud platforms across high-profile government institutions, Middle Eastern retail conglomerates, and fast-paced tech startups.
Beyond code and cloud platforms, I am an active contributor to open-source software, spending the last 4+ years sharing deep technical knowledge as a writer for the Fedora Project.
Professional Background
Senior DevOps Engineer
Architecting and owning the cloud infrastructure for an enterprise data governance SaaS platform, delivering secure, highly available deployments aligned with SDAIA & NDMO framework specifications to high-profile government and institutional clients across Saudi Arabia.
Terraform, ensuring scalable, repeatable, and auditable environments aligned with SDAIA (Saudi Data & AI Authority) and NDMO (National Data Management Office) regulations.Jenkins as a central automation server and engineered modular Ansible playbooks, eliminating manual ops toil and standardizing multi-client configuration management.Senior DevOps Engineer
Contracted to modernize DevOps practices at one of the Middle East's largest retail conglomerates, owning mobile release pipelines and cloud infrastructure automation end-to-end.
Fastlane, automating full build, test, and release workflows to App Store and Google Play — significantly reducing release friction.Terraform and Ansible, guaranteeing consistency across development, staging, and production environments.Helm charts, standardizing rollouts and eliminating manual intervention in production.Senior DevOps Engineer
Owned the complete DevOps lifecycle at a fast-scaling social media platform, building the infrastructure foundation to deploy code from commit to production with high velocity and zero downtime.
Jenkins, reducing manual deployment overhead and enabling low-risk microservice releases.Python, Bash, and Ansible that accelerated deployment cycles and eliminated human error.Zabbix for infrastructure metrics and ELK Stack for MongoDB cluster log analysis for proactive incident detection.Site Reliability Engineer
DevOps Engineer
Ansible for system & service configuration management.课外活动与社区
Technical Writer
Authoring and publishing deep-dive technical articles on container runtime engines (Podman), Linux system administration, and CLI workflows for the global Fedora developer community.
工作原则 (Manual of Me)
这是关于我如何沟通、协作和解决问题的个人说明书,旨在帮助团队更高效地配合。
文档优先 & 异步沟通
我深信文档的力量。我倾向于记录架构设计决策(ADR)、SRE 流程和排障手册(Runbooks),以减少团队间的重复沟通阻力,并支持高效的异步工作协作。
拒绝重复劳动与智能自动化
如果一件事情需要手动做两次,它就应该被编写成自动化脚本或声明式 IaC(Terraform/Ansible/Helm)。智能自动化是减少运维阻力的关键。
只设置可执行的有效告警
无效的报警是运维的大忌。我只关注并设置直接影响用户体验、可被排查和具备明确操作手册的报警,努力消除团队成员的“报警疲劳”。
无指责事后分析与持续学习
系统故障是学习的最佳契机。我推崇在团队内倡导无指责(Blameless)的事后回顾,剖析根本原因,寻找系统设计上的缺陷而非责备个人。