DevOps / SRE Engineer
Production infrastructure · Kubernetes · Observability · CI/CD · IaC
DevOps / SRE engineer with 4+ years operating production infrastructure. I run fleets of bare-metal and cloud servers, build observability and CI/CD from scratch, and codify everything as IaC. I care about uptime, fast safe deploys, and infrastructure you can reason about.
- 🌏 Based in Indonesia
- ☸️ Day-to-day: GCP/GKE, Kubernetes, Ansible, Prometheus/Grafana, Docker, GitHub Actions
- 🧰 I write tooling in Python · Go · Bash
- 🔧 On-call, incident response, postmortems — SRE practices, not just pipelines
Infrastructure & operations
- Operate 60+ production servers across GCP (managed Kubernetes — GKE) and bare-metal in three European DCs (Hetzner, Vultr, OVH); migrated a large share of workloads from bare-metal into GKE.
- Kubernetes in production: Helm charts, sealed-secrets, cert-manager, NetworkPolicies, resource quotas, dev/staging/prod environments, per-namespace service isolation.
Observability (built from zero)
- Stack: Prometheus · Grafana · Loki · Promtail · Alertmanager · Alerta — 100% production alert coverage routed to Telegram.
- Custom Prometheus exporters in Python and Go.
CI/CD & automation
- CI/CD on GitHub Actions, GitLab CI, Semaphore CI with zero-downtime deploys — deploy time cut from ~30 min to ~3 min.
- Multi-stage Docker builds, image optimization, publishing to GHCR.
- Infrastructure as Code with Ansible (roles, playbooks, dynamic inventory) — new environment spin-up from ~2 h to ~15 min.
Networking & security
- Nginx (reverse proxy, TLS, rate limiting), cert-manager for automated certificates, Cloudflare (WAF, DDoS protection, DNS, caching).
- Linux hardening: ufw/iptables, fail2ban, SSH policy (keys + jumphost), sealed-secrets, a security baseline on every node.
Data & internal tooling
- PostgreSQL & SQLite in production — deploys, schemas, backups, analytical SQL for operational metrics.
- Built an internal Server Dashboard — real-time monitoring UI for a fleet of 90+ servers (Prometheus + Ansible inventory + service registry, live over WebSocket).
SRE practices
- On-call rotation, incident response, postmortems; deploy standards and code review for infrastructure code.
A sample of public work — CI/CD pipelines and node/validator tooling I can show end-to-end:
drmed_ai— full GitHub Actions CI/CD:lint → typecheck → build+ semver-tagged Docker releases to GHCRcosmos-voting-bot— governance-proposal monitor, deployed via GitHub Actions over SSH with PM2cosmos-exporter— Prometheus metrics for Cosmos validators/wallets (Go)sui_voter— SUI gas-price auto-voter daemon (Python)
Web3 infrastructure experiments (Story Protocol)
story-subgraph— The Graph subgraph indexing IP assets, license terms, derivative links & royaltiesstory-ip-graph-mcp— MCP server for IP-graph ops, verified on-chain onaeneidstory-ip-explorer— Next.js dashboard for ecosystem metrics + IP lineagestory-ip-agent-demo— autonomous agent registering derivative IP on-chain


