Skip to content
// SYSTEM OPERATIONAL · AIOPS & MLOPS ARCHITECT

Hi, I'm Yazan.

I'm an AIOps Engineer. I specialize in fusing AI Agent Engineering with MLOps to transform infrastructure, turning manual operations into intelligent, self-healing systems powered by SRE and DevOps best practices.

Over the past 6+ years, I've architected and scaled cloud platforms across high-profile government institutions, Middle Eastern retail conglomerates, and fast-paced tech startups.

Beyond code and cloud platforms, I am an active contributor to open-source software, spending the last 4+ years sharing deep technical knowledge as a writer for the Fedora Project.

Experience 6+ Years
Core Focus AIOps / SRE
Open Source Fedora Writer
Security Air-Gapped
// CAREER TELEMETRY

Professional Background

[ ACTIVE // CURRENT ROLE ] Feb 2026 — Present
LOC // SAUDI ARABIA · ON-SITE

Senior DevOps Engineer

Governata

Architecting and owning the cloud infrastructure for an enterprise data governance SaaS platform, delivering secure, highly available deployments aligned with SDAIA & NDMO framework specifications to high-profile government and institutional clients across Saudi Arabia.

>_
Cloud Infrastructure & Governance: Architected and provisioned production-grade AWS infrastructure using Terraform, ensuring scalable, repeatable, and auditable environments aligned with SDAIA (Saudi Data & AI Authority) and NDMO (National Data Management Office) regulations.
>_
Air-Gapped & Government Deployments: Designed and deployed fully isolated, air-gapped environments for tier-1 government clients, including the Royal Court of Saudi Arabia, the Ministry of Health (MOH) , the General Entertainment Authority (GEA) , and integrations serving the Unified National Platform (my.gov.sa) , meeting strict national security and data sovereignty requirements.
>_
Automation & Operations: Implemented Jenkins as a central automation server and engineered modular Ansible playbooks, eliminating manual ops toil and standardizing multi-client configuration management.
Stack // Terraform AWS Ansible Jenkins Air-Gapped SDAIA & NDMO Compliance IaC
Contract Sep 2024 — Dec 2025
LOC // AMMAN, JORDAN · REMOTE

Senior DevOps Engineer

Majid Al Futtaim

Contracted to modernize DevOps practices at one of the Middle East's largest retail conglomerates, owning mobile release pipelines and cloud infrastructure automation end-to-end.

>_
Mobile CI/CD Pipelines: Designed and implemented CI/CD pipelines for Android and iOS engineering teams using Fastlane, automating full build, test, and release workflows to App Store and Google Play — significantly reducing release friction.
>_
Infrastructure as Code: Managed and maintained multi-environment cloud infrastructure using Terraform and Ansible, guaranteeing consistency across development, staging, and production environments.
>_
Kubernetes Deployments: Automated containerized microservice deployments to AWS EKS clusters using Helm charts, standardizing rollouts and eliminating manual intervention in production.
Stack // Terraform Ansible AWS EKS Helm Kubernetes CI/CD Mobile DevOps
Full-Time May 2022 — Jul 2024
LOC // AMMAN, JORDAN · ON-SITE

Senior DevOps Engineer

Baaz, Inc.

Owned the complete DevOps lifecycle at a fast-scaling social media platform, building the infrastructure foundation to deploy code from commit to production with high velocity and zero downtime.

>_
CI/CD Architecture: Designed end-to-end CI/CD pipelines using Jenkins, reducing manual deployment overhead and enabling low-risk microservice releases.
>_
Automation at Scale: Built automated operations scripts using Python, Bash, and Ansible that accelerated deployment cycles and eliminated human error.
>_
Observability Stack: Deployed a unified telemetry setup combining Zabbix for infrastructure metrics and ELK Stack for MongoDB cluster log analysis for proactive incident detection.
>_
ChatOps & SRE On-Call: Engineered real-time alerting using AWS Chatbot, SNS, and CloudWatch directly to Slack. Acted as primary on-call SRE lead for high-severity production incidents.
Stack // Jenkins Ansible Python AWS ELK Stack Zabbix
Full-Time Oct 2021 — May 2022
LOC // AMMAN, JORDAN · HYBRID

Site Reliability Engineer

Vardot

>_
Backup & Restore Automation: Built automated Jenkins backup and disaster recovery jobs from scratch to safeguard high-traffic enterprise digital platforms.
>_
Multi-Cloud Operations: Managed and monitored infrastructure across AWS, OVH, DigitalOcean, and Platform.sh, optimizing for uptime and performance.
>_
Incident Response: On-call SRE handling production triage, rapid resolution, and blameless post-mortem writing.
Stack // Jenkins AWS OVH DigitalOcean Linux
Full-Time Nov 2020 — Jun 2021
LOC // AMMAN, JORDAN · ON-SITE

DevOps Engineer

NADSOFT

>_
Configured CI/CD automation and utilized Ansible for system & service configuration management.
>_
Implemented server monitoring using Nagios across DigitalOcean and AWS infrastructure.
Stack // Ansible Nagios DigitalOcean AWS
// COMPLETE CAREER VERIFICATION & RECOMMENDATIONS AVAILABLE ON LINKEDIN linkedin.com/in/yazanmonshed
// 开源与社区

课外活动与社区

Contributor Nov 2020 — Present

Technical Writer

Fedora Project (Writer)

Authoring and publishing deep-dive technical articles on container runtime engines (Podman), Linux system administration, and CLI workflows for the global Fedora developer community.

Podman Linux Containers Fedora Magazine System Admin
// OPERATIONAL PHILOSOPHY

工作原则 (Manual of Me)

这是关于我如何沟通、协作和解决问题的个人说明书,旨在帮助团队更高效地配合。

文档优先 & 异步沟通

我深信文档的力量。我倾向于记录架构设计决策(ADR)、SRE 流程和排障手册(Runbooks),以减少团队间的重复沟通阻力,并支持高效的异步工作协作。

// PRINCIPLE 01

拒绝重复劳动与智能自动化

如果一件事情需要手动做两次,它就应该被编写成自动化脚本或声明式 IaC(Terraform/Ansible/Helm)。智能自动化是减少运维阻力的关键。

// PRINCIPLE 02

只设置可执行的有效告警

无效的报警是运维的大忌。我只关注并设置直接影响用户体验、可被排查和具备明确操作手册的报警,努力消除团队成员的“报警疲劳”。

// PRINCIPLE 03

无指责事后分析与持续学习

系统故障是学习的最佳契机。我推崇在团队内倡导无指责(Blameless)的事后回顾,剖析根本原因,寻找系统设计上的缺陷而非责备个人。

// PRINCIPLE 04
// CONNECT

Elsewhere

让我们携手合作

有项目想法?让我们讨论如何协助您完成 AIOps、AI 智能体工程、MLOps 或云基础架构。

可以通过 GitHubXLinkedIn 联系我。

A.I.D.A. SYSTEM MONITOR

Uplink status: SECURE
Agent role: DevOps Assistant