Why Work at Lenovo
Description and Requirements
*Please Note* This is a hybrid role in Morrisville, NC. This candidate will be required to work onsite three days a week.
This candidate MUST be a US citizen or US national; US permanent residents or candidates requiring sponsorship cannot be considered.
Position Overview
The Staff Researcher in AI Compute and Data Infrastructure will conduct applied research and hands-on development for intelligent, efficient, and resilient Hybrid AI systems. This position works across AI algorithms, computer systems, distributed computing, and data infrastructure to address performance, scalability, reliability, and energy-efficiency challenges.
The successful candidate will independently own substantial research and development workstreams, build production-quality software, characterize AI workloads, diagnose infrastructure issues, and develop cross-layer optimization technologies spanning GPUs and other accelerators, CPUs, memory, storage, networking, system software, data pipelines, and AI frameworks.
Key Responsibilities
- Research and develop technologies for AI compute and data infrastructure, distributed AI systems, and intelligent infrastructure management.
- Design and implement production-quality software, system components, services, APIs, diagnostic tools, and scalable data-processing pipelines.
- Characterize AI training, inference, and data-processing workloads using profiling, tracing, benchmarking, telemetry, logs, and hardware performance counters.
- Diagnose performance bottlenecks and reliability issues across GPUs, accelerators, CPUs, memory hierarchy, storage, networking, operating systems, runtimes, and AI frameworks.
- Develop hardware/software co-optimization solutions for GPU utilization, workload scheduling, resource allocation, memory and cache management, communication, data movement, storage access, and model execution.
- Optimize large-scale data ingestion, preprocessing, transformation, storage, retrieval, and delivery for AI training, inference, and analytics workloads.
- Build intelligent infrastructure diagnostics for anomaly detection, root-cause analysis, performance regression detection, system health assessment, capacity forecasting, and predictive maintenance.
- Develop fault-tolerance and resilience mechanisms, including fault detection and isolation, checkpointing, recovery, retry, failover, graceful degradation, and automated remediation.
- Apply machine learning and deep learning to workload modeling, performance prediction, resource optimization, failure prediction, and operational decision-making.
- Apply time-series analysis and signal processing to infrastructure telemetry, event detection, change-point detection, workload forecasting, and system health monitoring.
- Apply causal inference to performance attribution, root-cause analysis, intervention evaluation, and infrastructure optimization.
- Develop knowledge graphs to model infrastructure topology, hardware/software dependencies, workloads, operational events, and failure relationships.
- Optimize systems for throughput, latency, scalability, availability, resource utilization, energy consumption, and total cost of ownership.
- Collaborate with hardware, systems, software, architecture, and product teams to transition research technologies into Enterprise AI and Personal AI products.
- Contribute to patents, invention disclosures, technical publications, internal reports, and reusable software assets.
- Provide technical guidance and mentorship to junior researchers and engineers.
Minimum Qualifications
- Bachelor's degree in computer science, computer engineering, artificial intelligence, electrical engineering, applied mathematics, or a related field, or equivalent practical experience.
- Three or more years of relevant experience in AI compute and data infrastructure, machine learning systems, distributed systems, data platforms, performance engineering, reliability engineering, or advanced software development.
- Strong programming skills in Python, C++, Java, Go, Rust, Scala, or a comparable language.
- Demonstrated ability to design, implement, test, debug, profile, and optimize reliable software systems.
- Experience with system profiling, telemetry analytics, observability, performance diagnosis, or failure analysis.
- Technical expertise in at least two of the following areas:
- Machine learning or deep learning
- GPU or accelerator optimization
- Distributed training or inference systems
- Large-scale data processing
- Hardware/software co-optimization
- Time-series modeling or signal processing
- Infrastructure reliability and fault tolerance
- Causal inference
- Knowledge graphs or graph machine learning
- Strong analytical, experimental, communication, and cross-functional collaboration skills.
Preferred Qualifications
- Experience with PyTorch, TensorFlow, JAX, CUDA, ROCm, Spark, Flink, Ray, Kafka, Kubernetes, or related technologies.
- Experience with cloud, edge, on-premises, or hybrid AI infrastructure.
- Experience delivering research prototypes or advanced software into production environments.
- Publications, patents, open-source contributions, or demonstrated product impact.