Lead Site Reliability Engineer
工作概要:
The Lead Site Reliability Engineer is a seasoned subject matter expert who drives reliability, scalability, and performance for critical Disney Experiences platforms that power immersive guest interactions across theme parks, resorts, cruise, vacation, travel, retail, and consumer experiences. In this lead role, you will guide other engineers, set technical direction for complex systems, and ensure our digital and physical experiences remain highly available, secure, and resilient for guests around the world.
Responsibilities:
As a Lead Site Reliability Engineer, you will serve as a skilled, experienced problem-solver and team player, acting as the "go-to" technical lead for assigned platforms and services across Disney Experiences. You will architect and evolve cloud-native and hybrid infrastructures, champion observability and DevOps practices, and mentor individual contributors to deliver reliable, secure, and cost‑effective solutions that support Disney's creative, customer-focused, and innovative guest experiences. This role matters because it safeguards the technology behind our stories, requiring deep technical expertise, strong collaboration, and the ability to explain complex concepts and influence diverse stakeholders.
- Architect, design, and build scalable, maintainable, and secure infrastructure and platforms, including cloud-native and container-based solutions, to support mission-critical commerce and guest-facing applications.
- Lead the evolution of DevOps and SRE practices by consulting on, designing, and supporting CI/CD pipelines, automating infrastructure and operations, and creating telemetry and observability for monitoring and incident response.
- Serve as the SRE subject matter expert and technical lead for assigned products and platforms, owning reliability strategies, defining SLIs/SLOs/SLAs, and driving continuous improvement in uptime and performance.
- Identify root causes of operational issues in large-scale distributed systems, lead major incident response, and deliver clear retrospectives and remediation plans that reduce future risk and operational toil.
- Develop, maintain, and enhance automation, scripts, and Infrastructure as Code to standardize deployments, improve reliability, and support complex, non-standard environments without relying solely on runbooks.
- Collaborate with product, engineering, security, and operations teams to plan capacity, monitoring, configuration, security, metrics, reporting, recovery, and migration strategies for new initiatives and events impacting supported platforms.
- Mentor, train, and guide other engineers by providing continuous coaching, feedback, and technical direction, holding self and others accountable to commitments and aligning team work with organizational goals.
- Plan and coordinate team efforts and platform-oriented projects with moderate complexity and risk, breaking down organizational goals into clear outcomes and negotiating solutions to complex reliability challenges.
- Apply FinOps and cost-optimization principles to cloud environments, implementing governance, tagging, rightsizing, and usage analysis to balance reliability, performance, and cost efficiency.
- Champion a diverse, inclusive, team-oriented culture that encourages innovation, creative problem solving, and service-minded collaboration, ensuring every voice is heard and Disney values are experienced daily.
Required Qualifications:
- Minimum 7 years of related work experience in Site Reliability Engineering, Systems Engineering, or software development, with a focus on large-scale, distributed, and cloud-based systems.
- Bachelor's Degree in Computer Science, Information Systems, Engineering, or a related technical field, or equivalent work experience.
- Extensive hands-on experience with cloud hosting services (AWS, Azure, Google Cloud) and modern cloud architectures, including containers and orchestration platforms such as Docker, Kubernetes, ECS, AKS, and GKE.
- Proficiency in Infrastructure as Code and configuration management tools (e.g., Terraform, CloudFormation, Ansible, Chef) and CI/CD pipelines using tools such as GitHub, GitLab, Jenkins, AWS CodeBuild, or Azure DevOps.
- Fluency in core scripting and programming languages (e.g., Python, NodeJS, Golang, Bash, Perl, Ruby, Java) and strong UNIX/Linux administration, troubleshooting, and security skills.
- Applied expertise in observability and monitoring, including defining and implementing SLIs, SLOs, SLAs and using major APM and logging tools (e.g., AppDynamics, New Relic, ELK stack, Datadog, Splunk, New Relic).
- Strong knowledge of networking and distributed systems, including HTTP, TCP/IP, DNS, TLS, SSH, VPCs, gateways, firewalls, and microservices architectures.
- Experience with databases and data platforms such as MySQL, MongoDB, DynamoDB, Redis, and data solutions like Snowflake or Tableau, including ELT processes for data-driven decision making.
- Demonstrated ability to lead technical projects, evaluate new systems and infrastructure solutions for feasibility, and design reliable, scalable enterprise systems in agile environments.
- Outstanding troubleshooting methodology and communication skills, with the ability to explain difficult concepts, influence without direct authority, and mentor, train, and guide other engineers.
- Ability to plan and prioritize work aligned with organizational goals, collaborate effectively across teams, and hold self and others accountable to meet commitments.
Preferred Qualifications:
- Experience leveraging AI and automation for predictive insights and continuous improvement in system reliability and operational efficiency.
- Expertise in cloud infrastructure design and dynamic development technologies using Java, NodeJS, Python, and relational databases in large-scale business environments.
- Background operating production container environments and multi-origin hybrid (cloud and on-premise) architectures.
- Master's degree in Computer Science, Information Systems, or a related technical discipline.
關於Disney Experiences:
Disney Experiences 透過世界各地的主題公園、度假村、郵輪、獨特的度假體驗、產品等,將迪士尼故事和特許經營權的魔力帶入生活。迪士尼在旅遊業中大放異彩,在美國、歐洲和亞洲擁有六大度假勝地;一流的郵輪航線;廣受歡迎的度假擁有權計劃;以及屢獲殊榮的家庭探險導遊業務。此外,迪士尼的全球消費品業務還包括全球領先的授權業務、全球最大的兒童出版品牌、全球最大的跨平台遊戲授權商之一,以及遍佈全球和網絡的迪士尼商店。
關於 The Walt Disney Company:
Walt Disney Company 連同其子公司和聯營公司,是領先的多元化國際家庭娛樂和媒體企業,其業務主要涉及三個範疇:Disney Entertainment、ESPN 及 Disney Experiences。Disney 在 1920 年代的起步之初,只是一間卡通工作室,至今已成為娛樂界的翹楚,並昂然堅守傳承,繼續為家庭中每位成員創造世界一流的故事與體驗。Disney 的故事、人物與體驗傳遍世界每個角落,深入人心。我們在 40 多個國家/地區營運業務,僱員及演藝人員攜手協力,創造全球和當地人們都珍愛的娛樂體驗。
這個職位隸屬於 Disney (India) Private Limited,其所屬的業務部門是 Disney Experiences。
欲了解更多關於 Disney 針對應徵者使用 AI 的政策 請按此。
遇到技術問題?查看常見問題以尋求協助。
招聘流程
-
您的故事從哪裡開始?
探索 Disney 職位空缺和 The Life at Disney 網誌,了解華特迪士尼公司有待發掘的所有精彩機會。
-
迪士尼的故事裏,有你更精彩成就迪士尼故事
有許多不同品牌和業務可供探索。當您找到適合您的機會後,請填寫您的申請,進行下一步。
-
下一章
申請後,您將收到一封電子郵件,讓您可存取應徵者控制面板。建立您的登入資料,並確保經常檢視您的控制面板,以查看申請進度。
探索此地點 印度
The Walt Disney Company 運用精采故事的非凡力量,為世界各地獻上頂級娛樂、豐富資訊及靈感啟發,締造出使我們成為全球頂尖娛樂公司的知名品牌、創意理念及創新科技。
相關工作
我們的文化
相關內容
-
Career Development為何如此多的人選擇在迪士尼開啟職業生涯?三大理由 -
-
登記收取職缺通知
即時收到最新的工作機會的資訊。
分享
連結會在新分頁中開啟。