Alexandra Chen
DevOps Engineer
Profile
Senior DevOps Engineer with 8+ years designing resilient CI/CD pipelines, cloud-native platforms, and AI-assisted delivery workflows. Expert in Kubernetes, infrastructure as code, and production observability. Proven track record reducing deployment cycles from weeks to minutes while maintaining enterprise security and compliance standards.
Experience
Senior DevOps Engineer
NebulaStream Technologies - San Francisco, CALead platform engineering for real-time data streaming products serving 500K+ daily active users. Architect multi-region Kubernetes infrastructure on AWS and GCP. Own CI/CD pipeline strategy, SLO definitions, and incident response protocols. Mentor team of four DevOps engineers and drive infrastructure cost optimization initiatives.
Global Streaming Platform Resilience
- Rearchitected monolithic deployment pipeline into microservices-based GitOps workflow for a high-throughput streaming data platform processing 2M events per second. Replaced manual deployment processes with fully automated ArgoCD-based continuous delivery.
- AWS EKS across us-east-1 and eu-west-1, ArgoCD for GitOps, Terraform for IaC, Prometheus/Grafana for observability, Flux for secret management, custom admission controllers for policy enforcement
- Led infrastructure redesign from concept to production. Defined SLOs and error budgets. Built custom operators for canary analysis. Established disaster recovery procedures with automated failover.
- Tech Stack: Kubernetes ArgoCD Terraform Prometheus AWS EKS GitOps
- Challenges: 99.99% Availability Achievement: Reduced critical incidents by 73% through automated canary deployments and real-time rollback capabilities based on custom latency and error-rate metrics.
- Challenges: Deployment Velocity: Decreased mean lead time for changes from 14 days to 45 minutes through pipeline parallelization and artifact promotion strategies.
- Challenges: Cost Optimization: Reduced cloud infrastructure spend by $2.3M annually through rightsizing, spot instance adoption, and automated resource scheduling.
- Challenges: AI-Assisted Incident Response: Integrated Claude Code for automated log analysis and root-cause hypothesis generation during incidents, reducing MTTR from 90 minutes to 22 minutes.
Compliance Automation Framework
- Built automated compliance validation pipeline for SOC2 and GDPR requirements across 15 microservices. Integrated policy-as-code with human-in-the-loop approval gates.
- Open Policy Agent, Kyverno, Tekton, Vault, custom webhook services
- Designed policy framework and validation pipeline. Integrated with existing CI/CD. Built self-service compliance dashboard for engineering teams.
- Tech Stack: Open Policy Agent Kyverno HashiCorp Vault Tekton Policy as Code
- Challenges: Audit Readiness: Reduced compliance audit preparation from 6 weeks to 3 days through continuous automated evidence collection.
Junior DevOps Engineer
DataForge Systems - Seattle, WAMaintained and improved CI/CD infrastructure for analytics platform team. Supported migration from on-premises data centers to AWS. Automated routine operational tasks and contributed to monitoring and alerting improvements.
AWS Migration and Containerization
- Migrated 12 legacy applications from VM-based deployment to Docker containers on Amazon ECS. Established initial CI/CD pipelines and infrastructure automation practices.
- Amazon ECS, Docker, Jenkins, CloudFormation, CloudWatch, AWS CodeCommit
- Executed containerization of legacy applications. Built CloudFormation templates. Configured CloudWatch dashboards and alarms. Documented operational runbooks.
- Tech Stack: Amazon ECS Docker Jenkins CloudFormation AWS
- Challenges: Migration Completion: Completed migration 3 weeks ahead of schedule with zero data loss incidents.
- Challenges: Operational Efficiency: Reduced environment provisioning time from 3 days to 30 minutes through infrastructure automation.
Monitoring and Alerting Standardization
- Unified fragmented monitoring tools into cohesive observability stack with actionable alerting.
- Prometheus, Grafana, PagerDuty, custom exporters
- Standardized metric collection. Configured alert routing and escalation. Built team-specific dashboards.
- Tech Stack: Prometheus Grafana PagerDuty Observability
- Challenges: Alert Fatigue Reduction: Reduced false-positive alerts by 68% through alert tuning and dependency-aware suppression rules.
Personal Info
- San Francisco, CA
- al***@email.com
- 141****0147
- March 15, 1988