V2X logo

AI-Enabled Manager (Monitoring & Workload)

V2X
Remote, USARemoteAI/MLSenior-Level$140,000 - $200,000Posted: 4 days ago
Apply NowOpens V2X's site

About the role

63083

United States

Remote

Information Technology

US

Current Openings

Job Description

Overview

About Us

Working across the globe, V2X builds smart solutions designed to integrate physical and digital infrastructure from base to battlefield. We bring 120 years of successful mission support to improve security, streamline logistics, and enhance readiness. Aligned around a shared purpose, our $4.5B company and 16,000 people work alongside our clients, here and abroad, to tackle their most complex challenges with integrity, respect, responsibility, and professionalism.

Responsibilities

What You'll Do:

The AI-Enabled Manager, Monitoring & Workload Health is a technical leadership role responsible for the continuous observability, performance, and resilience of the enterprise IT ecosystem. Reporting to the Director of Adaptive Infrastructure Network & Security, this leader will drive the transition from reactive infrastructure monitoring to proactive, self-healing operations.

By leveraging agentic AI alongside enterprise monitoring platforms such as Microsoft System Center Operations Manager (SCOM) and SolarWinds Orion, the Manager will optimize the health of cloud workloads, on-premises infrastructure, network devices, and SaaS applications. This role is critical to establishing a resilient AIOps environment where predictive failure analysis and automated remediations eliminate operational downtime and reduce manual toil for the engineering teams.

Key Responsibilities:

AIOps Strategy & Advanced Monitoring

Design, deploy, and manage enterprise monitoring solutions, specifically focusing on optimizing SolarWinds Orion and Microsoft SCOM across a hybrid IT environment.

Integrate network telemetry, system logs, and application performance data into a centralized AIOps platform to support intelligent operations and anomaly detection.

Define and enforce monitoring baselines, thresholds, and intelligent alerting rules that prioritize business-critical SaaS applications and AI workloads.

Develop observability strategies that provide end-to-end visibility into complex infrastructure paths, including Zero Trust network segments and multi-cloud environments.

Predictive Failure & Agentic Remediation

Implement agentic AI workflows to identify degradation patterns and execute predictive failure analysis before system outages occur.

Build and maintain automated remediation runbooks utilizing tools like Ansible, Terraform, and Python to allow AI agents to execute self-healing actions on infrastructure and network devices.

Establish strict governance and human-in-the-loop oversight thresholds for agent-driven automated remediations, ensuring safe and compliant execution within the production environment.

Lead post-incident reviews (RCA) leveraging AI-generated timelines and telemetry data to continuously tune predictive algorithms and prevent recurring issues.

Workload Health & Performance Optimization

Monitor AI model inference traffic, LLM API calls, and agent-to-agent communication, ensuring the underlying infrastructure meets required Quality of Service (QoS) and latency SLAs.

Collaborate with infrastructure engineering and cloud teams to right-size compute workloads based on automated capacity planning and performance trend analysis.

Develop and deliver performance-based KPIs to IT leadership, highlighting system uptime, mean time to remediate (MTTR), and the effectiveness of automated resolution rates.

Team Leadership & Continuous Improvement

Manage the day-to-day tasking, performance, and effectiveness of a team of monitoring engineers and AIOps developers.

Foster a culture of automation-first thinking, guiding the team to codify operational policies and build "Compliance-as-Code" and "Self-healing" network capabilities.

Manage enterprise vendor and supplier agreements related to monitoring platforms, ensuring technical debt is minimized and platforms are optimized for value.

Qualifications

Minimum Requirements

Education:

Bachelor’s degree in Information Technology, Computer Science, Systems Engineering, or a related field; OR an equivalent combination of education and experience from which comparable knowledge and job skills can be obtained. (One year related experience may be substituted for one year of education, if degree is required).

Experience:

5-10 years of experience designing and operating complex enterprise monitoring, infrastructure, or reliability engineering solutions.

Leadership:

Proven experience as a manager or team leader overseeing technical personnel and large-scale automation projects.

Other Requirements:

U.S. Citizenship

Required Skills & Competencies

Deep technical expertise in administering and scaling Microsoft SCOM and SolarWinds Orion in enterprise environments.

Strong proficiency in network automation, Infrastructure-as-Code (IaC), and scripting languages (PowerShell, Python, YAML, JSON) to drive automated remediation workflows.

Experience integrating AIOps tools and agentic AI models with ITSM platforms (e.g., ServiceNow) for closed-loop incident resolution.

Understanding of modern cloud architectures (Azure, AWS), virtualization (VMware, Hyper-V), and Zero Trust networking principles.

Exceptional collaboration skills to partner with service desk, security, and application teams to build a cohesive, automated operational fabric.

What We Bring

At V2X we strive to be market competitive in our total reward offerings.

The successful candidate’s starting pay will be based on, but not limited to, their job related skills, experience, qualifications, work location, and market conditions.

The following salary range is intended to display the value of the company’s base pay compensation and may be modified at the discretion of the company.

USD $ 140,000 - 200,000

Provided salary range minimum and maximum values correspond to variances between regional/geographic locations across the United States.

Please speak with a recruiter for additional information.

Employee benefits include the following:

Healthcare coverage

Retirement plan

Life insurance, AD&D, and disability benefits

Wellness programs

Paid time off, including holidays

Learning and Development resources

Employee assistance resources

Pay and benefits are subject to change at any time and may be modified at the discretion of the company, consistent with the terms of any applicable compensation or benefit plans.

At V2X, we are deeply committed to both equal employment opportunity, including protection for Veterans and individuals with disabilities, and fostering an inclusive and diverse workplace. We ensure all individuals are treated with fairness, respect, and dignity, recognizing the strength that comes from a workforce rich in diverse experiences, perspectives, and skills.

This commitment, aligned with our core Vision and Values of Integrity, Respect, and Responsibility, allows us to leverage differences, encourage innovation, and expand our success in the global marketplace, ultimately enabling us to best serve our clients.

Minimum requirements

  • Bachelor’s degree in IT, Computer Science, Systems Engineering, or equivalent experience; 5-10 years in enterprise monitoring or reliability engineering.
  • Proven leadership managing technical teams and large-scale automation projects.
  • U.S. Citizenship required.

This listing was parsed by AI and may not be complete. Check the official posting on V2X's site for the most accurate information.

Ready to apply?
Applications are handled on V2X's own careers site.
Apply Now

More roles at V2X

Senior Project Manager
Remote, USA · Remote
View →
3D Graphics & Animation Specialist – Technical Training
Remote, USA · Remote
View →
Supply Chain Analyst
Remote, USA · Remote
View →
AI-Ready Infrastructure Automation Engineer
Remote, USA · Remote
View →
AI-Ready Endpoint & MDM Engineer
Remote, USA · Remote
View →