你好,我是 Yazan
我是一名 AIOps 工程师,专注于将 AI Agent 工程与 MLOps 深度融合,利用 SRE 与 DevOps 最佳实践将传统运维转化为自愈自治的智能系统。
在过去的 6+ 年里,我为中东顶尖政府机构、大型零售集团以及快速成长的科技初创公司设计并扩展了高可用云端平台。
除了云架构与代码,我积极投身开源社区,作为 Fedora Project 的技术作者,4 年来持续分享深度技术见解。
工作经历
Senior DevOps Engineer
Architecting enterprise data governance SaaS cloud platform for top government clients in Saudi Arabia, complying with SDAIA & NDMO security and regulatory frameworks.
- Cloud Infrastructure & Governance: Provisioned production AWS environments using
Terraformaligned with SDAIA and NDMO compliance standards. - Air-Gapped Government Deployments: Engineered isolated air-gapped environments for tier-1 institutions including the Royal Court of Saudi Arabia, Ministry of Health (MOH), and General Entertainment Authority (GEA).
- Automation & Ops: Implemented
Jenkinspipelines and modularAnsibleplaybooks for configuration management.
Senior DevOps Engineer
Modernized DevOps & mobile deployment pipelines across Middle East retail platforms.
- Mobile CI/CD: Built iOS/Android automation with
Fastlanefor automated Store releases. - IaC & Kubernetes: Provisioned AWS EKS microservices using
TerraformandHelm.
Senior DevOps Engineer
Owned end-to-end DevOps and SRE incident response for a fast-scaling social media platform.
- CI/CD Pipelines: Automated releases via Jenkins, Python, and Ansible.
- Observability: Deployed Zabbix and ELK Stack for MongoDB cluster metrics and log parsing.
- On-Call SRE: Led incident response and ChatOps integrations via AWS Chatbot & Slack.
Site Reliability Engineer
Managed disaster recovery and multi-cloud infrastructure across AWS, OVH, and DigitalOcean.
DevOps Engineer
Configured CI/CD automation and Nagios server monitoring for Linux environments.
开源与社区贡献
Fedora Project (Writer)
为全球 Fedora 开发者社区撰写并发表关于 Podman、Linux 容器引擎及 CLI 工作流的深度技术文章。
工作原则 (Manual of Me)
关于我如何沟通、协作与应对工程挑战的核心指南。
文档优先 & 异步沟通
编写架构决策记录 (ADRs) 与 SRE 操作手册 (Runbooks),降低团队沟通门槛,赋能高效异步协作。
消除重复劳动与智能自动化
如果某项任务手动被执行了两次,即刻将其转化为代码与声明式 IaC 自动化配置。
高价值告警与精准监控
告警应直接关联真实用户体验并附带明确 Operation Runbook,从源头杜绝告警疲劳。
无指责事后总结 (Blameless Post-Mortems)
将故障作为持续改善与学习的契机,深入分析系统设计与流程机制缺陷而非责备个人。