DL
Kafka Admin_India
Diverse Lynx India
Location not stated
2 months ago
- Apache Kafka
- Confluent
- Kafka Connect
- mTLS
- Kerberos
- ACLS
- RBAC
- Vault
- VPC
- SOX
- GDPR
- HIPAA
- Vulnerability Management
- Prometheus
- Grafana
- OpenSearch
- OpenTelemetry
- JVM
- Terraform
- Helm
- GitOps
- JDBC
- Debezium
- Elasticsearch
- Flink
- SQL
- DNS
- TLS
2 months ago
Job Description:
- Provision, configure, and manage Kafka clusters (Apache Kafka Food and Beverage or ZooKeeper-backed if legacy), Confluent Platform components (Schema Registry, Kafka Connect, ksqlDB, REST Proxy, Control Center), and Confluent Cloud resources.
- Ensure high availability (HA) and resilience across multi‐AZ/region deployments; manage partition replication, min.insync.replicas (ISR), and rack awareness.
- Implement and maintain client quotas, broker quotas, throttling, and back‐pressure strategies.
- Plan and execute upgrades/patching with zero/minimal downtime; maintain version currency against EOL and security advisories.
- Own backup/restore workflows for metadata (e.g., cluster configs), Schema Registry, and Connect configs
- Implement authentication (mTLS/SASL: SCRAM, OAUTHBEARER with IdP, or Kerberos if legacy) and authorization (Kafka ACLs, Confluent RBAC).
- Enforce encryption in transit and at rest, secret management (Vault/AKV/SSM), secure endpoint exposure, and private connectivity (VPC peering/PrivateLink).
- Govern topic naming conventions, retention policies, schema evolution rules (backward/fully compatible), and PII handling.
- Maintain audit trails (broker, Schema Registry, Control Center, Confluent Cloud audit logs) and compliance artifacts (SOX, GDPR, HIPAA/PCI as applicable).
- Coordinate vulnerability management (CVE watch, hardening baselines, CIS/STIG alignment).
- Deploy and maintain metrics pipelines (JMX → Prometheus/Grafana), logs (ELK/OPensearch), and tracing (OpenTelemetry where applicable)
- Define and monitor SLOs/SLIs (produce/consume latency, end‐to‐end lag, broker CPU/heap/file handles, network throughput, GC pauses).
- Analyze broker hotspots, partition skew, large message handling, compaction and retention settings, and tune GC/JVM for sustained throughput.
- Implement consumer lag monitoring and alerting with actionable runbooks to prevent data loss or SLA breaches.
- Model workload growth, topic/partition scaling, storage/IOPS, network egress/ingress, and broker sizing.
- For Confluent Cloud, optimize environment → cluster → topic hierarchy, throughput tiers, retention costs, and data flow egress; right‐size Dedicated vs. Standard clusters.
- Design partitioning strategies to meet throughput targets while balancing consumer concurrency and avoiding small partitions anti‐patterns
- Automate cluster and topic lifecycle using Terraform (providers for Kafka/Confluent), Helm, GitOps workflows.
- Operate and scale Kafka Connect with distributed workers; manage connectors (JDBC, Debezium CDC, S3/ADLS/GCS, Elasticsearch, etc.), transforms (SMTs), and converter settings.
- Administer Schema Registry (compatibility levels, subject naming strategies, serializers/deserializers, schema size & references).
- Operate ksqlDB and/or Flink for streaming SQL; enforce resource limits and multi‐tenant governance.
- Engineer low‐latency and secure connectivity across data centers/VPCs/VNETs; manage DNS, TLS SANs, LB configurations, advertised.listeners, and inter‐cluster networking.
Kafka Admin_India · Diverse Lynx India