The Pega System Administrator is responsible for administering, supporting and operating Pega Platform 25.x in a client-managed, containerised Azure Kubernetes Service environment. The role provides hands-on platform administration, Helm-based deployment support, AKS runtime operations, secure configuration management, external service integration, production observability, performance tuning, incident diagnosis, patching and upgrade support across DEV, TEST, UAT and PROD environments.
The role requires strong capability across Pega platform configuration, Kubernetes orchestration, Azure networking and ingress, secrets and certificate management, SRS, Kafka, Elasticsearch/OpenSearch and related external service dependencies.
Key Accountabilities and main responsibilities
Strategic Focus
- Maintain and support a secure, scalable and highly available Pega Platform 25.x architecture on Azure Kubernetes Service.
- Contribute to platform strategy for operating Pega as Kubernetes pods, services, ingress, config maps and tiered workloads.
- Support correct workload segregation across WebUser, Background Processing, Search, BIX, ADM, Batch, RealTime, RTDG, Custom1 and Custom2 node types.
- Ensure Pega platform strategy appropriately accounts for required external services such as Kafka, SRS and Elasticsearch/OpenSearch, with Cassandra where decisioning workloads require it.
- Support high availability, scalability, controlled rollout, rollback and release readiness for Pega workloads
Operational Management
- Administer Pega Platform 25.x environments, including configuration, startup validation, node health and release readiness.
- Maintain and update Helm chart configuration including values.yaml, provider settings, deployment actions, tier definitions, JDBC parameters, image references, image pull secrets and security context settings.
- Troubleshoot pod restarts, failed rollouts, crash loops, failed mounts, container startup issues, readiness/liveness probes, services, namespaces, secrets and config maps.
- Operate and troubleshoot Kafka, SRS, Elasticsearch/OpenSearch and, where applicable, Cassandra dependencies.
- Manage and troubleshoot tiered workloads across web, batch, analytics, dataflow, search and decisioning-related node types.
- Support relational database connectivity, split-schema architecture, JDBC configuration, driver compatibility and connectivity troubleshooting.
- Use PDC or equivalent observability tools, Pega logs, alert logs and AKS events to triage incidents and identify root cause.
- Support install, deploy, upgrade, upgrade-deploy, patching, rollback and controlled release activities.
People Leadership
- Collaborate with application, platform, infrastructure, security, architecture and vendor teams to resolve platform incidents.