
Skills - AI & Agentic Systems
- Agentic AI Development (MCP)
- Large Language Models (LLM)
- Retrieval-Augmented Generation (RAG)
- Prompt Engineering
- Generative AI
- AI Governance
Skills - Cloud Infrastructure
- AWS Cloud Architecture and Services
- Google Cloud Platform (GCP)
- Infrastructure as Code (IaC)
- High Availability System Design
- Disaster Recovery Planning
Skills - DevOps & SRE
- Kubernetes Administration
- Docker Containerization
- CI/CD Pipeline Implementation
- Prometheus & Grafana Monitoring
- Jenkins Automation
- GitOps Practices
Skills - Programming & Development
- Golang (Backend Development)
- Python
- Shell Scripting
- RESTful API Design
- Microservices Architecture
- Version Control (Git)
Skills - Security & Compliance
- Cloud Security Architecture
- Financial Services Compliance
- Encryption Implementation
- Security Best Practices
- Access Management
Skills - System Design
- Distributed Systems
- Event Streaming (Kafka / Flink)
- Workflow Orchestration (Temporal.io)
- Database Design
- System Scalability
- Load Balancing & Caching
Skills - Leadership & Communication
- Enterprise Architecture (TOGAF)
- Technical Writing
- Team Mentorship
- Architecture Planning
- Cross-functional Collaboration
- Knowledge Sharing
$$
\Huge \textbf {Rezky Aulia Pratama} \\
\small {Solution Architect \; | \; Cloud-Native \cdot SRE \cdot Agentic AI}
$$
<aside>
<img src="notion://custom_emoji/f0c2c741-90c1-460a-b1d8-55fabbd9bee3/15d0f42d-7035-80b9-8c27-007a5f3cb030" alt="notion://custom_emoji/f0c2c741-90c1-460a-b1d8-55fabbd9bee3/15d0f42d-7035-80b9-8c27-007a5f3cb030" width="40px" /> Linkedin Profile
</aside>
<aside>
✉️ [email protected]
</aside>
<aside>
<img src="notion://custom_emoji/f0c2c741-90c1-460a-b1d8-55fabbd9bee3/15d0f42d-7035-80e8-978e-007ae44ef4f1" alt="notion://custom_emoji/f0c2c741-90c1-460a-b1d8-55fabbd9bee3/15d0f42d-7035-80e8-978e-007ae44ef4f1" width="40px" /> Medium
</aside>
<aside>
<img src="https://prod-files-secure.s3.us-west-2.amazonaws.com/f0c2c741-90c1-460a-b1d8-55fabbd9bee3/cca60680-50ba-4aa1-9b01-3edf949109e5/github.png" alt="https://prod-files-secure.s3.us-west-2.amazonaws.com/f0c2c741-90c1-460a-b1d8-55fabbd9bee3/cca60680-50ba-4aa1-9b01-3edf949109e5/github.png" width="40px" /> Github
</aside>
Profile
Cloud-native infrastructure was my foundation. Site Reliability Engineering hardened my instincts. Today I build the next layer — agentic AI systems that make financial infrastructure intelligent, resilient, and self-aware. Across 11+ years I have engineered distributed systems where failure is not an option: payments that clear in milliseconds, near real-time fraud detection, and reliability designed from the ground up rather than bolted on. As Solution Architect at Bank Sinarmas I lead enterprise architecture for Digital Banking, Branch Banking, and Middleware Systems. Earlier, I hardened crypto-exchange infrastructure at Pintu and built payment, settlement, and mobile systems at DANA Indonesia. My work sits at the intersection of cloud architecture, SRE, and applied AI — Go microservices on Kubernetes (GKE/EKS), Temporal-orchestrated financial workflows, and high-throughput Kafka/Flink pipelines.
Technical Excellence
Applied AI & Agentic Systems — 2026 Focus
- Designing LLM-powered agents on the MCP protocol — a Fraud Detection Agent and an Ops Copilot targeting production banking
- Building RAG pipelines from scratch (chunker, embedder, vector store, retriever, generator) over a banking-domain knowledge base
- Integrating multi-provider LLMs (AWS Bedrock, DeepSeek) into IT-resilience and automation workflows
- Authoring the “GenAI in Practice” article series with companion open-source code
Cloud Architecture Achievements
- Designed and implemented high-availability systems achieving 99.99% uptime
- Optimized cloud resources to reduce infrastructure costs by 40%
- Built scalable architectures processing millions of daily transactions
- Deployed critical financial services with zero downtime
Site Reliability Engineering Impact
- Implemented enterprise-wide monitoring using Prometheus and Grafana
- Automated incident response to reduce Mean Time To Recovery (MTTR) by 60%
- Enhanced deployment reliability by 85% through GitOps practices
- Developed disaster recovery protocols with RPO < 15 minutes
Recent Work Experience