Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com
C

Staff, Site Reliability Engineer (Tech Infra)

Coupang
๐Ÿ‡ฐ๐Ÿ‡ท South Korea
On-site
Staff / Principal
2 months ago
  • Incident Management
  • Disaster Recovery
  • Load Testing
  • Unix
  • Linux
  • Python
  • Java
  • Golang
  • Ruby
  • TCP/IP
  • AWS
  • Azure
  • GCP
  • CI/CD
  • IaC
  • Devops
  • Terraform
  • Docker
  • Kubernetes
  • Prometheus
  • Grafana
  • Elastic Stack
  • Datadog
  • New Relic
  • JVM
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

ํšŒ์‚ฌ ์†Œ๊ฐœ

์ฟ ํŒก์€ ๊ณ ๊ฐ ๊ฐ๋™ ์‹คํ˜„์„ ์œ„ํ•ด ์กด์žฌํ•ฉ๋‹ˆ๋‹ค. ๊ณ ๊ฐ๋“ค์ด "์ฟ ํŒก ์—†์ด ๊ทธ๋™์•ˆ ์–ด๋–ป๊ฒŒ ์‚ด์•˜์„๊นŒ?" ๋ผ๊ณ  ๋งํ•  ๋•Œ, ๋น„๋กœ์†Œ ์šฐ๋ฆฌ์˜ ๋ฏธ์…˜์„ ์‹คํ˜„ํ•˜๊ณ  ์žˆ์Œ์„ ์•Œ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค. ๊ณ ๊ฐ๋“ค์˜ ์‡ผํ•‘๊ณผ ์‹์‚ฌ, ์ƒํ™œ ์ „๋ฐ˜์„ ํŽธํ•˜๊ฒŒ ๋งŒ๋“ค๊ฒ ๋‹ค๋Š” ์œ ์ผํ•œ ์ง‘๋…์œผ๋กœ ์ฟ ํŒก์€ ์ˆ˜์–ต ๋‹ฌ๋Ÿฌ ๊ทœ๋ชจ์˜ ์ด์ปค๋จธ์Šค ์‚ฐ์—… ์ „๋ฐ˜์˜ ํ˜์‹ ์„ ์ด๋Œ๊ณ  ์žˆ์Šต๋‹ˆ๋‹ค. ์ฟ ํŒก์€ ๊ฐ€์žฅ ๋น ๋ฅด๊ฒŒ ์„ฑ์žฅํ•˜๋Š” ์ด์ปค๋จธ์Šค ๊ธฐ์—… ์ค‘ ํ•˜๋‚˜๋กœ, ๊ตญ๋‚ด ์ปค๋จธ์Šค ์—…๊ณ„์—์„œ์˜ ๋…๋ณด์ ์ธ ์ž…์ง€์™€, ๊ณ ๊ฐ ์‹ ๋ขฐ๋ฅผ ๊ตฌ์ถ•ํ–ˆ์Šต๋‹ˆ๋‹ค.ย 

์ฟ ํŒก์€ ์Šคํƒ€ํŠธ์—… ๋ฌธํ™”๋ฅผ ๊ธฐ๋ฐ˜์œผ๋กœ ํ•œ ๊ธ€๋กœ๋ฒŒ ๋Œ€ํ˜• ์ƒ์žฅ์‚ฌ๋ผ๊ณ  ์ž๋ถ€ํ•ฉ๋‹ˆ๋‹ค. ์ด๊ฒƒ์ด ์ฐฝ๋ฆฝ ๋‹น์‹œ์˜ ๊ธฐ๋ฏผํ•จ์„ ์ง€ํ•˜๋ฉฐ, ์‹ ๊ทœ ์„œ๋น„์Šค๋ฅผ ๋Š์ž„์—†์ด ์ถœ์‹œํ•˜๋ฉฐ ๋น„์ฆˆ๋‹ˆ์Šค๋ฅผ ํ™•์žฅํ•ด ๋‚˜๊ฐ€๋Š” ์šฐ๋ฆฌ์˜ ์„ฑ์žฅ ๋™๋ ฅ์ž…๋‹ˆ๋‹ค. ์ฟ ํŒก์˜ ๋ชจ๋“  ์ž„์ง์›์—๊ฒŒ๋Š” ๊ธฐ์—…๊ฐ€ ์ •์‹ ์„ ๊ฐ–์ถ”๊ณ  ์ƒˆ๋กœ์šด ํ˜์‹ ๊ณผ ์ด๋‹ˆ์…”ํ‹ฐ๋ธŒ๋ฅผ ์ถ”์ง„ํ•  ์ˆ˜ ์žˆ๋Š” ๊ธฐํšŒ๊ฐ€ ์ฃผ์–ด์ง‘๋‹ˆ๋‹ค. ์ฃผ์ € ์—†์ด ์ผ์— ๋›ฐ์–ด๋“ค์–ด ์„ฑ๊ณผ๋ฅผ ์ด๋ฃจ๊ณ ์ž ํ•˜๋Š” ๊ณผ๊ฐ์„ฑ์ด, ๋ฐ”๋กœ ์ฟ ํŒก์ด ์ผํ•˜๋Š” ๋ฐฉ์‹์˜ ๋ณธ์งˆ์ž…๋‹ˆ๋‹ค. ์ฟ ํŒก์—์„œ๋Š” ์—ฌ๋Ÿฌ๋ถ„ ์ž์‹ , ๋™๋ฃŒ, ํŒ€ ๊ทธ๋ฆฌ๊ณ  ํšŒ์‚ฌ ์ „์ฒด๊ฐ€ ๋งค์ผ ์„ฑ์žฅํ•˜๋Š” ๋ชจ์Šต์„ ๋ชฉ๊ฒฉํ•  ๊ฒƒ์ž…๋‹ˆ๋‹ค.ย 

์ฟ ํŒก์˜ ๋ชจ๋“  ์ง์›์€ ์ปค๋จธ์Šค์˜ ๋ฏธ๋ž˜๋ฅผ ๋งŒ๋“ค๊ฒ ๋‹ค๋Š” ์ฟ ํŒก์˜ ๋ฏธ์…˜์— ์ง„์‹ฌ์ž…๋‹ˆ๋‹ค. ์šฐ๋ฆฌ๋Š” ๊ณ ๊ฐ์˜ ๋ฌธ์ œ๋ฅผ ํ•ด๊ฒฐํ•ด ๋‚˜๊ฐ€๊ณ , ์ „ํ†ต์ ์ธ ๊ด€๋…๊ณผ ํ†ต๋…์— ๋งž์„œ๋ฉฐ ์‹คํ˜„ ๊ฐ€๋Šฅํ•œ ํ•œ๊ณ„๋ฅผ ๋›ฐ์–ด๋„˜๊ณ  ์žˆ์Šต๋‹ˆ๋‹ค. ๊ณ ๊ฐ€์šฉ์„ฑ (always-on) ๊ณผ ์ตœ์ฒจ๋‹จ์˜ ์•ž์„  ๊ธฐ์ˆ  (high-tech), ์ดˆ์—ฐ๊ฒฐ์‚ฌํšŒ (hyper-connected world) ์—์„œ์˜ ๋†€๋ผ์šด ์—…๋ฌด ๊ฒฝํ—˜์„ ์›ํ•˜์‹ ๋‹ค๋ฉด, ์ง€๊ธˆ ๋ฐ”๋กœ ์ฟ ํŒก์— ํ•ฉ๋ฅ˜ํ•˜์„ธ์š”.

ย 

์ง๋ฌด ์†Œ๊ฐœ

์ฟ ํŒก์˜ Site Reliability Engineer(SRE)๋Š” ์†Œํ”„ํŠธ์›จ์–ด ์—”์ง€๋‹ˆ์–ด๋ง๊ณผ ์‹œ์Šคํ…œ ์—”์ง€๋‹ˆ์–ด๋ง์„ ๊ฒฐํ•ฉํ•˜์—ฌ, ๋Œ€๊ทœ๋ชจ ์ด์ปค๋จธ์Šค ์‹œ์Šคํ…œ์„ ๊ตฌ์ถ•ยท์šด์˜ยทํ™•์žฅํ•˜๋Š” ํ•ต์‹ฌ์ ์ธ ์—ญํ• ์ž…๋‹ˆ๋‹ค.SRE ํŒ€์˜ ์ผ์›์œผ๋กœ์„œ, ๋ชจ๋“  ๊ณ ๊ฐ facing ์„œ๋น„์Šค๊ฐ€ ์•ˆ์ •์ ์œผ๋กœ ์šด์˜๋˜๊ณ , ์ง€์†์ ์œผ๋กœ ๋ชจ๋‹ˆํ„ฐ๋ง๋˜๋ฉฐ, ์ž๋™ํ™”๋˜์–ด ์žˆ๊ณ , ํ™•์žฅ ๊ฐ€๋Šฅํ•˜๊ฒŒ ์„ค๊ณ„๋˜๋„๋ก ์ฑ…์ž„์ง€๊ฒŒ ๋ฉ๋‹ˆ๋‹ค. SRE ์กฐ์ง์€ โ€˜์šด์˜์„ ์—”์ง€๋‹ˆ์–ด๋ง ๋ฌธ์ œ๋กœ ํ•ด๊ฒฐํ•œ๋‹คโ€™๋Š” ์ฒ ํ•™ ์•„๋ž˜, ์ž๋™ํ™”๋ฅผ ์ตœ์šฐ์„ ์œผ๋กœ ์ ‘๊ทผํ•ฉ๋‹ˆ๋‹ค.์ด ํฌ์ง€์…˜์—์„œ๋Š” Observability, Incident Management, Disaster Recovery, Load Testing, Capacity Engineering ๋“ฑ ๋‹ค์–‘ํ•œ ์˜์—ญ์—์„œ ์ตœ๊ณ  ์ˆ˜์ค€์˜ ์ธํ”„๋ผ ์ž๋™ํ™”๋ฅผ ๊ตฌ์ถ•ํ•˜๊ฒŒ ๋ฉ๋‹ˆ๋‹ค. ๋˜ํ•œ ์ œํ’ˆ ๊ฐœ๋ฐœ ์ดˆ๊ธฐ ๋‹จ๊ณ„๋ถ€ํ„ฐ ์ฐธ์—ฌํ•˜์—ฌ, ์‹ค์ œ ์šด์˜ ์ค‘ ๋ฐœ์ƒํ•˜๋Š” ์ด์Šˆ ํ•ด๊ฒฐ๊นŒ์ง€ ์ „ ๊ณผ์ •์— ๊ฑธ์ณ ๊ฐœ๋ฐœ ์กฐ์ง๊ณผ ๊ธด๋ฐ€ํžˆ ํ˜‘์—…ํ•ฉ๋‹ˆ๋‹ค.๋”๋ถˆ์–ด ์šด์˜ ์„œ๋น„์Šค์˜ SLI/SLA ๊ธฐ์ค€์„ ์œ ์ง€ํ•˜๊ณ , SRE ์›์น™๊ณผ ๋ฒ ์ŠคํŠธ ํ”„๋ž™ํ‹ฐ์Šค๋ฅผ ์กฐ์ง ์ „๋ฐ˜์— ํ™•์‚ฐ์‹œํ‚ค๋Š” ์—ญํ• ๋„ ์ˆ˜ํ–‰ํ•ฉ๋‹ˆ๋‹ค.๋Œ€๊ทœ๋ชจ ๋ถ„์‚ฐ ์‹œ์Šคํ…œ ํ™˜๊ฒฝ์—์„œ ๋ณต์žกํ•œ ๊ธฐ์ˆ  ๋ฌธ์ œ ํ•ด๊ฒฐ์— ์—ด์ •์ด ์žˆ๊ณ , ๋†’์€ ์ฑ…์ž„๊ฐ์„ ๊ฐ€์ง€๊ณ  ํŒ€ ๊ฐ„ ํ˜‘์—…๊ณผ ์ปค๋ฎค๋‹ˆ์ผ€์ด์…˜์„ ์›ํ™œํ•˜๊ฒŒ ์ˆ˜ํ–‰ํ•  ์ˆ˜ ์žˆ๋‹ค๋ฉด ์ง€๊ธˆ ์ฟ ํŒก์— ํ•ฉ๋ฅ˜ํ•˜์„ธ์š”!ย ย ์—…๋ฌด ๋‚ด์šฉ
  • ์ฟ ํŒก์˜ ๋ชจ๋“  ๊ณ ๊ฐ ๋Œ€์ƒ ์„œ๋น„์Šค์˜ ์•ˆ์ •์„ฑ, ์ƒํƒœ, ์„ฑ๋Šฅ์„ ์ฑ…์ž„์ง€๋Š” ์ฃผ์š” ๋‹ด๋‹น์ž๋กœ ์—ญํ•  ์ˆ˜ํ–‰
  • ์ฟ ํŒก ์• ํ”Œ๋ฆฌ์ผ€์ด์…˜์˜ ์›Œํฌํ”Œ๋กœ์šฐ์™€ ์˜์กด์„ฑ์— ๋Œ€ํ•œ ๊นŠ์€ ์ดํ•ด ํ™•๋ณด
  • ์‹œ์Šคํ…œ ๊ฐ€์šฉ์„ฑ, ์„ฑ๋Šฅ, ์•ˆ์ •์„ฑ๊ณผ ๊ด€๋ จ๋œ KPI ๋ฐ SLO ์ •์˜ ๋ฐ ๊ด€๋ฆฌ
  • ์‹ ์†ํ•œ ์žฅ์•  ๋ณต๊ตฌ, ์šด์˜ ๋ฆฌ๋ทฐ ๋ฐ ์‚ฌํ›„ ๋ถ„์„์„ ํฌํ•จํ•œ Incident Management ํ”„๋กœ์„ธ์Šค ๋ฐ ์ž๋™ํ™” ๊ตฌ์ถ•
  • ํšจ๊ณผ์ ์ธ ๋ชจ๋‹ˆํ„ฐ๋ง, ์•Œ๋ฆผ, ํ…”๋ ˆ๋ฉ”ํŠธ๋ฆฌ ์‹œ์Šคํ…œ ๊ตฌ์ถ• ๋ฐ ์šด์˜์„ ์œ„ํ•œ ๋ฒ ์ŠคํŠธ ํ”„๋ž™ํ‹ฐ์Šค ์ˆ˜๋ฆฝ
  • ์„œ๋น„์Šค ์„ฑ์žฅ์— ๋Œ€๋น„ํ•˜๊ธฐ ์œ„ํ•œ ์ •๊ธฐ์ ์ธ Disaster Recovery ํ…Œ์ŠคํŠธ ๋ฐ Load Testing ์ž๋™ํ™” ๊ตฌ์ถ•
  • ์ œํ’ˆ ๊ฐœ๋ฐœ ํŒ€๊ณผ ๊ธด๋ฐ€ํžˆ ํ˜‘๋ ฅํ•˜์—ฌ ํ™•์žฅ์„ฑ๊ณผ ์šด์˜ ์šฉ์ด์„ฑ์„ ๊ณ ๋ คํ•œ ์„ค๊ณ„ ๊ตฌํ˜„
  • ์„œ๋น„์Šค ์•ˆ์ •์„ฑ์„ ์œ ์ง€ํ•˜๊ธฐ ์œ„ํ•œ ํ”„๋กœ๋•์…˜ ๋ฐฐํฌ ๊ฐ€๋“œ๋ ˆ์ผ ๋ฐ ์ž๋™ํ™” ๊ตฌ์ถ•
  • 24x7 ์˜จ์ฝœ ๋กœํ…Œ์ด์…˜ ์ฐธ์—ฌ ๋ฐ ๋น ๋ฅธ ์†๋„์˜ ํ™˜๊ฒฝ์—์„œ ๋ฌธ์ œ ๋Œ€์‘
  • ์กฐ์ง ๋‚ด ๋‹ค์–‘ํ•œ ๋ ˆ๋ฒจ๊ณผ ํšจ๊ณผ์ ์œผ๋กœ ์ปค๋ฎค๋‹ˆ์ผ€์ด์…˜

ย 

์ž๊ฒฉ ์š”๊ฑด

  • ๋Œ€๊ทœ๋ชจ ๋ถ„์‚ฐ ์‹œ์Šคํ…œ ๊ตฌ์ถ• ๋ฐ ์šด์˜ ๊ฒฝ๋ ฅ 5๋…„ ์ด์ƒ
  • UNIX/Linux ์‹œ์Šคํ…œ์— ๋Œ€ํ•œ ๊นŠ์€ ์ดํ•ด์™€ ์šด์˜ ๊ฒฝํ—˜
  • Python, Java, Golang, Ruby ์ค‘ ํ•˜๋‚˜ ์ด์ƒ์˜ ํ”„๋กœ๊ทธ๋ž˜๋ฐ ์—ญ๋Ÿ‰
  • ์‹œ์Šคํ…œ, ๋„คํŠธ์›Œํฌ(TCP/IP), ์ฝ”๋“œ ์ „๋ฐ˜์— ๊ฑธ์นœ ๋ฌธ์ œ ํ•ด๊ฒฐ ๋ฐ ๋ถ„์„ ๋Šฅ๋ ฅ (๋ฐ์ดํ„ฐ ๊ธฐ๋ฐ˜ ์˜์‚ฌ๊ฒฐ์ • ํฌํ•จ)
  • AWS, Azure, Google Cloud Platform ๋“ฑ ํด๋ผ์šฐ๋“œ ์ธํ”„๋ผ ๊ฒฝํ—˜
  • CI/CD, IaC ๋“ฑ DevOps ๋ฐ SRE ๊ด€๋ จ ์‹ค๋ฌด ์ดํ•ด (Terraform ์‚ฌ์šฉ ๊ฒฝํ—˜ ์šฐ๋Œ€)
  • Docker, Kubernetes ๋“ฑ ์ปจํ…Œ์ด๋„ˆ ๋ฐ ์˜ค์ผ€์ŠคํŠธ๋ ˆ์ด์…˜ ๊ธฐ์ˆ  ๊ฒฝํ—˜
  • ๋‹ค์–‘ํ•œ ์กฐ์ง๊ณผ ๊ธฐ์ˆ  ์˜์—ญ ๊ฐ„ ํ˜‘์—…์ด ๊ฐ€๋Šฅํ•œ ์ปค๋ฎค๋‹ˆ์ผ€์ด์…˜ ์—ญ๋Ÿ‰
  • Prometheus, Grafana, Elastic Stack, Datadog, New Relic๋“ฑ Observability๋„๊ตฌ๊ฒฝํ—˜

ย 

์šฐ๋Œ€ ์‚ฌํ•ญ

  • ์ปดํ“จํ„ฐ๊ณตํ•™, ์—”์ง€๋‹ˆ์–ด๋ง ๋˜๋Š” ๊ด€๋ จ ๋ถ„์•ผ ํ•™์‚ฌ ํ•™์œ„
  • ๋Œ€๊ทœ๋ชจ ์›น ๊ธฐ๋ฐ˜ Java ์•„ํ‚คํ…์ฒ˜ ๋ฐ JVM ์„ค์ • ๊ฒฝํ—˜
  • ํด๋ผ์šฐ๋“œ, ๋ชจ๋‹ˆํ„ฐ๋ง ๋“ฑ ๊ด€๋ จ ๊ธฐ์ˆ  ์ž๊ฒฉ์ฆ ๋ณด์œ 
  • ๋Œ€๊ทœ๋ชจ ์ด์ปค๋จธ์Šค ํ”Œ๋žซํผ ๊ฒฝํ—˜

ย 

๊ทผ๋ฌด์ง€: ์ฟ ํŒก ์„ ๋ฆ‰ ์˜คํ”ผ์Šคย 

์ „ํ˜• ์ ˆ์ฐจ ๋ฐโ€ฏ์•ˆ๋‚ดโ€ฏ์‚ฌํ•ญย 

  • ์ „ํ˜•โ€ฏ์ ˆ์ฐจย 
    • ์„œ๋ฅ˜์ „ํ˜•(์˜๋ฌธ์ด๋ ฅ์„œ ์ œ์ถœ) -ย  ํ™”์ƒ๊ธฐ์ˆ ๋ฉด์ ‘ 1์ฐจ - ํ™”์ƒ๊ธฐ์ˆ ๋ฉด์ ‘ 2์ฐจย  โ€“ ์ตœ์ข… ํ•ฉ๊ฒฉ
    • ์ „ํ˜•์ ˆ์ฐจ๋Š” ์ง๋ฌด๋ณ„๋กœ ๋‹ค๋ฅด๊ฒŒ ์šด์˜๋  ์ˆ˜ ์žˆ์œผ๋ฉฐ, ์ผ์ • ๋ฐ ์ƒํ™ฉ์— ๋”ฐ๋ผ ๋ณ€๋™๋  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.
    • ์ „ํ˜• ์ผ์ • ๋ฐ ๊ฒฐ๊ณผ๋Š” ์ง€์›์„œ์— ๋“ฑ๋กํ•˜์‹  ์ด๋ฉ”์ผ๋กœ ๊ฐœ๋ณ„ ์•ˆ๋‚ด ๋“œ๋ฆฝ๋‹ˆ๋‹ค.ย 
  • ์ฐธ๊ณ โ€ฏ์‚ฌํ•ญย 
    • ๋ณธ ๊ณต๊ณ ๋Š” ๋ชจ์ง‘ ์™„๋ฃŒ ์‹œ ์กฐ๊ธฐ ๋งˆ๊ฐ๋  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.
    • ์ง€์›์„œ ๋‚ด์šฉ ์ค‘ ํ—ˆ์œ„์‚ฌ์‹ค์ด ์žˆ๋Š” ๊ฒฝ์šฐ์—๋Š” ํ•ฉ๊ฒฉ์ด ์ทจ์†Œ๋  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.
    • ์ทจ์—… ๋ณดํ˜ธ ๋Œ€์ƒ์ž(๋ณดํ›ˆ๋Œ€์ƒ์ž, ์žฅ์• ์ธ ๋“ฑ)๋Š” ๊ด€๋ จ ๋ฒ•๋ฅ ์— ๋”ฐ๋ผ ์ฑ„์šฉ์šฐ๋Œ€๋ฅผ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.
    • ์ง๊ธ‰๊ณผ ๋‹ด๋‹น ์—…๋ฌด ๋ฒ”์œ„๋Š” ํ›„๋ณด์ž์˜ ์ „๋ฐ˜์ ์ธ ๊ฒฝ๋ ฅ๊ณผ ๊ฒฝํ—˜ ๋“ฑ ์ œ๋ฐ˜์‚ฌ์ •์„ ๊ณ ๋ คํ•˜์—ฌ ๋ณ€๊ฒฝ๋  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค. ์ด๋Ÿฌํ•œ ๋ณ€๊ฒฝ์ด ํ•„์š”ํ•  ๊ฒฝ์šฐ, ์ตœ์ข… ํ•ฉ๊ฒฉ ํ†ต์ง€ ์ „ ์ ์ ˆํ•œ ์‹œ๊ธฐ์— ํ›„๋ณด์ž์™€ ์ปค๋ฎค๋‹ˆ์ผ€์ด์…˜ ๋  ์˜ˆ์ •์ž…๋‹ˆ๋‹ค.
    • ์ฑ„์šฉ ๋ฐ ์—…๋ฌด ์ˆ˜ํ–‰๊ณผ ๊ด€๋ จํ•˜์—ฌ ์š”๊ตฌ๋˜๋Š” ๋ฒ•๋ น์ƒ ์ž๊ฒฉ์ด ๊ฐ–์ถ”์–ด์ง€์ง€ ์•Š์€ ๊ฒฝ์šฐ ์ฑ„์šฉ์ด ์ œํ•œ๋  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.ย 
    • ๊ณ ์šฉํ˜•ํƒœ๋Š”ย ์ •๊ทœ์ง์œผ๋กœย ์ˆ˜์Šต๊ธฐ๊ฐ„ย 12์ฃผ๋ฅผย ํฌํ•จํ•ฉ๋‹ˆ๋‹ค. ๋‹จ, ์—…๋ฌด์ƒ ํ•„์š”ํ•œ ๊ฒฝ์šฐ์—๋Š” ์ƒ๊ธฐ ์ˆ˜์Šต๊ธฐ๊ฐ„์„ ์ ์šฉํ•˜์ง€ ์•Š๊ฑฐ๋‚˜ ๋‹จ์ถ• ๋˜๋Š” ์—ฐ์žฅํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.ย ย ย 

๊ฐœ์ธ์ •๋ณด ์ฒ˜๋ฆฌ๋ฐฉ์นจโ€ฏย ย 

  • ์ฟ ํŒก ๊ทธ๋ฃน์€ ์ž…์‚ฌ์ง€์›์ž ๊ฐœ์ธ์ •๋ณด ์ฒ˜๋ฆฌ๋ฐฉ์นจ(์•„๋ž˜ ๋งํฌ)์— ๋”ฐ๋ผ ๊ท€ํ•˜์˜ ๊ฐœ์ธ์ •๋ณด๋ฅผ ์ˆ˜์ง‘ํ•˜์—ฌ ์ฒ˜๋ฆฌํ•ฉ๋‹ˆ๋‹ค.โ€ฏhttps://www.coupang.jobs/kr/privacy-policyโ€ฏย ย 

์„œ๋ฅ˜โ€ฏ๋ฐ˜ํ™˜ ์ •์ฑ…โ€ฏย ย 

  1. ๋ณธย ๊ณ ์ง€๋Š” ใ€Ž์ฑ„์šฉ์ ˆ์ฐจ์˜๊ณต์ •ํ™”์—๊ด€ํ•œ๋ฒ•๋ฅ ใ€ย ์ œ11์กฐ์ œ6ํ•ญ์— ๋”ฐ๋ฅธ ๊ฒƒ ์ž…๋‹ˆ๋‹ค.ย 
  2. ๋‹น์‚ฌ ์ฑ„์šฉ์— ์‘์‹œํ•œ ๊ตฌ์ง์ž ์ค‘ ์ตœ์ข… ํ•ฉ๊ฒฉ์ด ๋˜์ง€ ๋ชปํ•œ ๊ตฌ์ง์ž๋Š” ใ€Ž์ฑ„์šฉ์ ˆ์ฐจ์˜ ๊ณต์ •ํ™”์— ๊ด€ํ•œ ๋ฒ•๋ฅ ใ€์— ๋”ฐ๋ผ ์ œ์ถœํ•œ ์ฑ„์šฉ์„œ๋ฅ˜์˜ ๋ฐ˜ํ™˜์„ ์ฒญ๊ตฌํ•  ์ˆ˜ ์žˆ์Œ์„ ์•Œ๋ ค ๋“œ๋ฆฝ๋‹ˆ๋‹ค. ๋‹ค๋งŒ, ํ™ˆํŽ˜์ด์ง€ ๋˜๋Š” ์ „์ž์šฐํŽธ์œผ๋กœ ์ œ์ถœ๋œ ๊ฒฝ์šฐ๋‚˜ ๊ตฌ์ง์ž๊ฐ€ ๋‹น์‚ฌ์˜ ์š”๊ตฌ ์—†์ด ์ž๋ฐœ์ ์œผ๋กœ ์ œ์ถœํ•œ ๊ฒฝ์šฐ์—๋Š” ๊ทธ๋Ÿฌํ•˜์ง€ ์•„๋‹ˆํ•˜๋ฉฐ, ์ฒœ์žฌ์ง€๋ณ€์ด๋‚˜ ๊ทธ ๋ฐ–์— ๋‹น์‚ฌ์—๊ฒŒ ์ฑ…์ž„ ์—†๋Š” ์‚ฌ์œ ๋กœ ์ฑ„์šฉ์„œ๋ฅ˜๊ฐ€ ๋ฉธ์‹ค๋œ ๊ฒฝ์šฐ์—๋Š” ๋ฐ˜ํ™˜ํ•œ ๊ฒƒ์œผ๋กœ ๋ด…๋‹ˆ๋‹ค.
  3. ์œ„2ํ•ญ ๋ณธ๋ฌธ์— ๋”ฐ๋ผ ์ฑ„์šฉ ์„œ๋ฅ˜ ๋ฐ˜ํ™˜ ์ฒญ๊ตฌ๋ฅผ ํ•˜๋Š” ๊ตฌ์ง์ž๋Š” ์ฑ„์šฉ ์„œ๋ฅ˜ ๋ฐ˜ํ™˜ ์ฒญ๊ตฌ์„œ [์ฑ„์šฉ์ ˆ์ฐจ์˜ ๊ณต์ •ํ™”์— ๊ด€ํ•œ ๋ฒ•๋ฅ  ์‹œํ–‰๊ทœ์น™ ๋ณ„์ง€ ์ œ 3 ํ˜ธ ์„œ์‹]๋ฅผ ์ž‘์„ฑํ•˜์—ฌ ์ด๋ฉ”์ผ (recruitingops@coupang.com) ๋กœ ์ œ์ถœํ•˜๋ฉด, ์ œ์ถœ์ด ํ™•์ธ๋œ ๋‚ ๋กœ๋ถ€ํ„ฐ 14 ์ผ ์ด๋‚ด์— ์ง€์ •ํ•œ ์ฃผ์†Œ์ง€๋กœ ๋“ฑ๊ธฐ์šฐํŽธ์„ ํ†ตํ•˜์—ฌ ๋ฐœ์†กํ•ด ๋“œ๋ฆฝ๋‹ˆ๋‹ค. ์ด ๊ฒฝ์šฐ ๋“ฑ๊ธฐ์šฐํŽธ์š”๊ธˆ์€ ์ˆ˜์‹ ์ž ๋ถ€๋‹ด์œผ๋กœ ํ•˜๊ฒŒ ๋˜์˜ค๋‹ˆ ์œ ๋…ํ•˜์‹œ๊ธฐ ๋ฐ”๋ž๋‹ˆ๋‹ค.โ€ฏย 
  4. ๋‹น์‚ฌ๋Š” ์œ„2ํ•ญ ๋ณธ๋ฌธ์— ๋”ฐ๋ฅธ ๊ตฌ์ง์ž์˜ ๋ฐ˜ํ™˜ ์ฒญ๊ตฌ์— ๋Œ€๋น„ํ•˜์—ฌ ์ฑ„์šฉ ์—ฌ๋ถ€๊ฐ€ ํ™•์ •๋œ ๋‚ ๋กœ๋ถ€ํ„ฐ 180 ์ผ๊ฐ„ ๊ตฌ์ง์ž๊ฐ€ ์ œ์ถœํ•œ ์ฑ„์šฉ์„œ๋ฅ˜ ์›๋ณธ์„ ๋ณด๊ด€ํ•˜๊ฒŒ ๋˜๋ฉฐ, ๊ทธ๋•Œ๊นŒ์ง€ ์ฑ„์šฉ์„œ๋ฅ˜์˜ ๋ฐ˜ํ™˜์„ ์ฒญ๊ตฌํ•˜์ง€ ์•„๋‹ˆํ•  ๊ฒฝ์šฐ์—๋Š” ใ€Ž๊ฐœ์ธ์ •๋ณด ๋ณดํ˜ธ๋ฒ•ใ€์— ๋”ฐ๋ผ ์ง€์ฒด ์—†์ด ์ฑ„์šฉ์„œ๋ฅ˜ ์ผ์ฒด๋ฅผ ํŒŒ๊ธฐํ•  ์˜ˆ์ •์ž…๋‹ˆ๋‹ค.
  5. ๋‹จ, ์œ„ 1ํ•ญ ๋‚ด์ง€ 4ํ•ญ์˜ ๋‚ด์šฉ์€ ๋Œ€ํ•œ๋ฏผ๊ตญ์˜ ๋…ธ๋™ ๊ด€๊ณ„ ๋ฒ•๋ น์ด ์ ์šฉ๋˜๋Š” ๊ฒฝ์šฐ์—๋งŒ ์ ์šฉ๋ฉ๋‹ˆ๋‹ค. ๊ทธ ์ด์™ธ์˜ ๊ฒฝ์šฐ์—๋Š” ์ ์šฉ๋˜์ง€ ์•Š์Šต๋‹ˆ๋‹ค.

About the Role:ย 

Site Reliability Engineers (SREs) at Coupang is a mission-critical role which combines software and system engineering to build, run and scale our complex, large-scale ecommerce systems. As part of the Site Reliability Engineering team, you will be responsible for ensuring all our customer facing services are healthy, monitored, automated, and designed to scale. As SRE organization we take pride in handling โ€œoperations as an engineeringโ€ problem with automation first approach. You will use your background to build best in class infrastructure automation for areas such as Observability, Incident management, Disaster Recovery, Load testing, Capacity engineering and many more. In this role you will work very closely with our product development teams from an early stage of design to all the way helping resolve any production incidents, maintaining SLI/SLA bar for production services and influencing them with SRE principles and best practices. If you take pride in complete ownership, have a passion for solving complex technical challenges for large scale distributed systems and demeanor to work and communicate effectively across team boundaries, this is the role for you!ย 

ย 

Key Responsibilities:

ยทย ย Serve as a primary point responsible for the reliability, health, and performance of all Coupang customer-facing services.ย 

ยทย ย Gain deep knowledge of Coupang application workflow and dependencies.ย 

ยทย ย Define and track key performance indicators (KPIs) and service-level objectives (SLOs) related to system availability, performance, and reliability.

ยทย ย Build world class incident management process and automation, including fast incident remediation, incident operational reviews and retrospectives.

ยทย ย Develop and implement best practices for creating and maintaining effective monitoring, alerting, and telemetry systems.

ยทย ย Build automation to execute regular Disaster Recovery testing and load testing to stay ahead of expected growth of Coupang services.ย 

ยทย ย Work closely with product development teams to ensure the products are designed with scale and operability in mind.ย 

ยทย ย Build right guardrails and automation for deploying production changes holding the reliability bar.ย 

ยทย ย Participate in a 24x7 rotation for production issue escalations, functions well in a fast-paced environment.ย 

ยทย ย Communicate effectively with people at all levels of the organization.

ย 

Essential Qualifications:

ยทย ย 5+ years of industry experience building and operating large scale distributed systems.ย 

ยทย ย Deep UNIX/Linux systems knowledge and administration background.

ยทย ย Demonstrated programming skills in one or more of: Python, Java, Golang, Ruby.

ยทย ย Strong problem-solving and analytical skills spanning systems, network (TCP/IP) and code, with a focus on data-driven decision-making.

ยทย ย Experience with cloud-based infrastructure, including AWS, Azure, or Google Cloud Platform.

ยทย ย Strong understanding of DevOps and SRE practices, including continuous integration, continuous delivery, and infrastructure as code (IaC). Experience with Terraform is a plus.

ยทย ย Experience with containerization and orchestration technologies, such as Docker and Kubernetes.

ยทย ย Excellent communication and collaboration skills, with the ability to work with teams across distinct functions and technical domains.

ยทย ย Knowledge of observability ecosystem including metrics, logging, tracing and tools, such as Prometheus, Grafana, Elastic Stack, Datadog, or New Relic.

ย 

Preferred Qualifications:

ยทย ย Bachelor's degree in computer science, engineering, or a related technical field.ย 

ยทย ย Prior experience working with large scale web-based Java architectures and JVM configuration.

ยทย ย Professional certifications in cloud platforms, monitoring tools, or related technologies.

ยทย ย Previous experience working on a large-scale eCommerce platform.ย 

ย 

Office:Seoul, Korea

Staff, Site Reliability Engineer (Tech Infra) ยท Coupang

Auto apply with Likeremote