I’m a Platform and Site Reliability Engineer, and I’ve spent 15+ years building the platforms other engineers build on. That work has been across AWS and Google Cloud, in finance, fintech, government, and media. The project that taught me the most was the AWS platform behind Colombia’s first fully digital bank. We ran it to a 99.999% availability SLO, with multi-region disaster recovery, self-service environments, and CI/CD that teams could actually trust. I also stood up a self-hosted LLM inference platform (vLLM/Ollama) so the bank could use models without regulated data ever leaving the building. More recently I moved a US startup’s production onto Google Cloud (GKE, Cloud Functions) in a gradual, zero-downtime cutover of about 25 services.
My bias is toward paved roads: reusable Terraform modules, GitOps delivery, and SLOs with error budgets, so product teams can move fast without babysitting infrastructure. I’m certified on Kubernetes and IaC (CKA, CKAD, Terraform Associate), but the certificates matter less than the habits behind them: on-call, incident response, blameless postmortems, real DR drills. I’ve led DevOps, SRE, and Cloud teams of up to 25. I’m based in Medellín, Colombia, and open to remote Platform Engineering, SRE, DevOps, or AI-infrastructure roles across US and EU hours, at Staff/Principal or senior IC scope.
What I work with day to day:
- Cloud & IaC: AWS (EC2, EKS/ECS, Lambda, RDS/Aurora, DynamoDB, S3), Google Cloud (GKE, Cloud Functions, API Gateway), Azure, Terraform, Pulumi, CloudFormation
- Platform & CI/CD: internal developer platforms, self-service provisioning, golden paths, reusable Terraform modules, environment promotion, Jenkins, GitHub Actions, GitLab CI, Argo CD (GitOps), canary and blue-green deployments
- Kubernetes & containers: Kubernetes, Docker, Helm, service mesh (Istio, Linkerd), Kubernetes RBAC
- Observability & reliability: Prometheus, Grafana, OpenTelemetry, Datadog, New Relic, ELK, SLOs and error budgets, incident management and RCA, disaster recovery and HA, FinOps
- Security & DevSecOps: SAST/DAST, OWASP Top 10, SIEM, WAF, vulnerability management (Nessus, Qualys, Dome9), IAM, zero-trust segmentation (Calico), mTLS, secrets management (Vault, AWS Secrets Manager, SOPS), SOC 2, PCI DSS, ISO 27001
- AI Infrastructure & LLM serving: self-hosted LLM inference platform (vLLM, Ollama on GPU-backed EC2), Amazon Bedrock, OpenAI-compatible model gateway and serving (FastAPI), hosted model APIs (OpenAI, Hugging Face, DeepSeek), OpenWebUI, LangChain RAG over logs and operational data, LLMOps/MLOps, AI platform engineering
- Data & databases: PostgreSQL, MySQL, Amazon Aurora, DynamoDB, Redshift, MongoDB, Apache Airflow, Spark, Amazon EMR, Kafka, Glue, DMS
- Languages & OS: Python, Go, Rust, Bash, Linux, Unix, FreeBSD
- Spoken: Spanish (native), English (professional working proficiency)