JACK SIU Résumé ↓

Platform · Site Reliability Engineering

Jack Siu

I build the infrastructure that everything else runs on — and the automation that keeps it quiet.

5+ years in production infrastructure at hyperscale cloud. Terraform and Kubernetes for the platform, Prometheus and Grafana for the signal, event-driven automation for the incidents nobody should have to page a human for.

Open to Platform / SRE / Infrastructure roles — NYC or remote

5+
Years running production infrastructure
~30m → <1m
Incident triage time, after automation
~10 hrs
On-call toil removed per week
60–70%
Reduction in developer onboarding time

Selected work

Systems I built and operated

The first one is the one worth reading — it ran in production for years and is the clearest picture of how I think about reliability.

Case study Multi-year in production

Incident Automation Engine

Event-driven system at Oracle Cloud Infrastructure that ingests alerts, system events, and failed health checks, then deduplicates, enriches, and drives the full incident lifecycle — create, route, resolve, notify — without a human in the loop. Structured telemetry streams to downstream consumers so other teams build on it instead of getting paged.

~30m → <1m
Triage MTTR per ticket
20+ / week
Incidents handled automatically
~10 hrs / week
On-call toil removed
PythonKubernetesTerraform Pub/SubGrafana
Read the architecture & tradeoffs

pebbl.

iOS and Android photo booth app — camera filters, sticker overlays, freemium subscriptions. Shipped to both stores from zero, with a CI/CD pipeline and RevenueCat monetization behind it.

React NativeExpoCI/CDRevenueCat
View project

DUPE

Pass-the-phone social deduction game for 3–8 players. Everyone gets a secret word except the Dupe, who gets a decoy. Custom SVG avatars, 110+ word packs, zero setup.

ReactSVGGame design
Play it

SmartMaze

Full-stack maze solver implementing BFS, A*, hill climbing, and genetic algorithms, racing them against each other in an interactive live demo.

Pythonp5.jsTkinter
Live demo

AI Property Manager

LLM-driven agent automating end-to-end property management — tenant communication, maintenance scheduling, and lease workflows against real-time integrations.

PythonReactLLMs
In progress

Experience

Where I've operated

Jun 2022 — Apr 2026
Oracle Cloud Infrastructure
Seattle, WA
Software Development Engineer II — Data Platform
  • Architected multi-region infrastructure on OCI with Terraform and Kubernetes, codifying provisioning across distributed data services.
  • Built the Prometheus and Grafana observability stack with SLOs, SLIs, and error-budget alerting that surfaced regressions before customers saw them — caught one that would have been a customer-facing outage on a major downstream consumer.
  • Automated incident triage with an event-driven service that creates, routes, and resolves tickets on its own — cut MTTR and ~10 hours/week of on-call toil. Led the blameless post-mortems that fed back into it.
  • Engineered real-time streaming pipelines emitting structured incident telemetry to internal consumers.
  • Cut developer onboarding time 60–70% with self-serve tooling; wrote the runbooks that were adopted org-wide.
  • Drove DevSecOps hardening: migrated hardcoded credentials into Vault, remediated CVEs, and patched dependencies across services.
2024 — Present
pebbl.
Remote
Founder & Lead Software Engineer
  • Shipped a cross-platform iOS/Android app from zero to production on the App Store and Google Play.
  • Designed the CI/CD pipelines and backend services behind real-time image processing and content delivery.
  • Built subscription monetization via RevenueCat, with in-app purchases and analytics.
Aug 2020 — May 2022
Trinnex
Philadelphia, PA
Backend Software Engineer
  • Built and shipped full-stack features (Angular/TypeScript, Node.js, Oracle) for secure government platforms under strict compliance requirements.
  • Engineered backend data pipelines and GeoJSON visualization services.
  • Optimized PL/SQL systems for query performance across compliance-bound datasets.

Stack

What I reach for

Cloud & infrastructure

TerraformKubernetesDocker HelmOCILinuxVault

Observability & reliability

PrometheusGrafanaSLOs / SLIs Error budgetsMTTRAlerting

SRE practice

Incident responseBlameless post-mortems On-callRunbooksDevSecOps

CI/CD & automation

GitOps / ArgoCDJenkinsIaC Git pipelinesMaven

Languages

PythonBashTypeScript Node.jsSQLGo (learning)

Contact

Let's talk reliability

Open to Platform, SRE, and Infrastructure roles — NYC or remote. Fastest way to reach me is email.