01

28–36 weeks

Production AI Engineer

Design, evaluate, deploy, and operate reliable AI systems.

1.1

ML & mathematical foundations

active

Build the mathematical and experimental judgment behind trustworthy models.

ProbabilityStatisticsOptimizationEvaluation
Proof: Classical ML study with leakage-safe evaluation and error analysis
4–5 weeks
1.2

Applied deep learning

active

Train, debug, evaluate, and serve a serious non-LLM PyTorch model.

PyTorchTraining loopsTransfer learningError analysis
Proof: Reproducible model, evaluation report, model card, and inference interface
5–6 weeks
1.3

LLM & generative AI engineering

next

Build evaluated RAG and constrained tool-using systems.

TransformersRAGAgentsEvals
Proof: Evaluated RAG system and bounded agent with failure analysis
5–6 weeks
1.4

Production ML engineering

locked

Operate versioned AI services with observability, rollout, and rollback.

MLOpsServingObservabilitySecurity
Proof: Load-tested, observable deployment with an automated quality gate
6–7 weeks
1.5

AI system design

locked

Make defensible data, model, scale, reliability, and cost decisions.

ArchitectureScaleReliabilityCost
Proof: Four reviewed AI system design documents
3–4 weeks
1.6

Large-company interviews

locked

Practice coding, ML depth, system design, and behavioral communication.

CodingML depthSystem designLeadership
Proof: Passing mock-interview rubric across all four interview types
5–7 weeks
02

28–40 weeks

Edge AI · NVIDIA

Export, optimize, profile, and operate accelerated AI on target hardware.

2.1

Systems programming

locked

Build the C++, Linux, memory, concurrency, and profiling foundation.

C++20LinuxMemoryConcurrency
Proof: Profiled native component with Python interoperability
5–6 weeks
2.2

Parallel computing & CUDA

locked

Reason about GPU execution, memory movement, and measured kernel performance.

CUDAMemoryStreamsNsight
Proof: Optimized kernels with an Nsight-based performance narrative
6–8 weeks
2.3

Model interchange & runtimes

locked

Move correct models from PyTorch through ONNX to accelerated runtimes.

ONNXGraphsExecution providersValidation
Proof: Correctness and performance matrix across CPU, CUDA, and TensorRT
4–5 weeks
2.4

TensorRT specialization

locked

Build and diagnose optimized engines across precision and dynamic shapes.

TensorRTQuantizationPluginsProfiling
Proof: FP16/INT8 engine report with quality and performance comparisons
6–7 weeks
2.5

Jetson & real-time AI

locked

Sustain a camera-to-inference pipeline under power and thermal constraints.

JetPackDeepStreamGStreamerThermals
Proof: Sustained real-time deployment with recovery and device measurements
5–7 weeks
2.6

Edge GenAI & multimodal

locked

Fit useful multimodal and language workflows within device budgets.

SLMsKV cacheLocal RAGRouting
Proof: Offline-first assistant with quality, memory, and latency evidence
4–5 weeks
2.7

Edge MLOps & fleet security

locked

Package, observe, update, verify, and recover deployed model fleets.

Fleet updatesTelemetryRollbackSecurity
Proof: Signed staged update with integrity, offline recovery, and rollback proof
4–5 weeks