PnP Pipeline On-call Support Engineer
- Role
- DevOps
The listing doesn't say where it hires from. Check the description or the employer's site before applying.
No BS summary
On-call support engineer for monitoring and triaging PnP/CI pipelines. Needs alert/dashboard monitoring experience, strong log troubleshooting, and ability to distinguish infrastructure, firmware, and analyzer failures.
Core skills
Optional skills
Role Summary Your primary mission is to ensure the consistent health and stability of our PnP (Power and Performance) pipelines. You will serve as the first line of defense, monitoring pipelines and performing initial triaging to reduce manual overhead for our engineering teams. This role is critical to maintaining the velocity of our post-silicon PnP work as we scale. Core Responsibilities Pipeline Monitoring: Conduct regular, proactive health checks on functional and PnP-enabled flows. Ensure data integrity by verifying that test results are populating correctly on the RLS PnP dashboard. Initial Triage: Investigate pipeline failures and categorize them to streamline resolution: Infrastructure/Deployment Issues: CI pools, hargow issues, or platform environments (e.g., Artemis/Coleman HW/FW mismatches). UTF / PnP Analyzer Issues: Troubleshoot problems within the Unified Test Framework (UTF) or artifact deployment. Mature Pipeline Issues: Address issues specific to pipeline logic for stable flows that are no longer under active development. Resolution & Escalation: Perform bug fixes for straightforward issues (e.g., patching deployment or execution changes). For infrastructure issues, post to the appropriate internal channels for escalation and follow-up. Escalate complex or systemic issues to pipeline owners by providing necessary logs and context. Collaborate with pipeline owners and XFN (Cross-Functional) DRIs to ensure proper placement of PnP trace tags and alignment of measurement windows. Qualifications & Technical Requirements Core Requirements: Proven experience in monitoring alerts and dashboards. Strong troubleshooting skills with the ability to identify patterns in failure logs and suggest improvements. Ability to differentiate between infrastructure, firmware, and analyzer-specific errors. Preferred Qualifications: Familiarity with jest-e2e and the Unified Test Framework (UTF). Experience working within internal infrastructure (e.g., buck, sandcastle, phabricator, internal AI tooling). Ability to read and interpret CI pipeline logs effectively. Ability to leverage AI tooling effectively (claude code, Metamate, etc.)
What you'll do
- Conduct regular, proactive health checks on functional and PnP-enabled flows.
- Ensure data integrity by verifying that test results are populating correctly on the RLS PnP dashboard.
- Investigate pipeline failures and categorize them to streamline resolution.
- Troubleshoot infrastructure/deployment issues such as CI pools, hargow issues, or platform environments, including Artemis/Coleman HW/FW mismatches.
- Troubleshoot problems within the Unified Test Framework (UTF) or artifact deployment.
- Address issues specific to pipeline logic for stable flows that are no longer under active development.
- Perform bug fixes for straightforward issues, such as patching deployment or execution changes.
- For infrastructure issues, post to the appropriate internal channels for escalation and follow-up.
- Escalate complex or systemic issues to pipeline owners by providing necessary logs and context.
- Collaborate with pipeline owners and XFN DRIs to ensure proper placement of PnP trace tags and alignment of measurement windows.
What they require
- Proven experience in monitoring alerts and dashboards.
- Strong troubleshooting skills with the ability to identify patterns in failure logs and suggest improvements.
- Ability to differentiate between infrastructure, firmware, and analyzer-specific errors.
- Preferred: Familiarity with jest-e2e and the Unified Test Framework (UTF).
- Preferred: Experience working within internal infrastructure, e.g. buck, sandcastle, phabricator, internal AI tooling.
- Preferred: Ability to read and interpret CI pipeline logs effectively.
- Preferred: Ability to leverage AI tooling effectively, such as claude code, Metamate, etc.