NVIDIA AI infrastructure and operations
Why it matters: it is the closest formal path from data-center operations to GPU clusters, CUDA, and inference serving.
NVIDIA workshopLearning
A working shelf, not a catalog dump. Each pick has a short note on why it matters for infrastructure and applied AI. I am not affiliated with or sponsored by the publishers listed here.
Editorial picks
Why it matters: it is the closest formal path from data-center operations to GPU clusters, CUDA, and inference serving.
NVIDIA workshopWhy it matters: tool use, memory, and multi-agent setups are the patterns behind the maintenance dashboard.
Start the courseWhy it matters: Martin Kleppmann’s reliability, scalability, and maintainability still describe the systems under the model layer.
Courses
Foundations, generative AI, cloud, and agent tooling. Longer blurbs stay folded.
Infrastructure and cloud
Why it matters: GPU operations, accelerated data science, and networking (Spectrum, BlueField) sit on the chips and infrastructure layers of the stack.
Why it matters: GPU vs CPU, on-prem and cloud placement, storage, networking, power, and cooling — the operator’s view of an AI cluster.
Stack: GPUs, NVIDIA, CUDA.
Why it matters: a cloud path for someone who already ships enterprise systems — Azure Machine Learning, cognitive services, and bot integration.
Stack: Azure, Python.
Why it matters: most of the lab apps only become real once they run as containers. Covers architecture, manifests, Minikube, kubectl, and a small web app with MongoDB.
Stack: Docker, Kubernetes, MongoDB, YAML.
Models, agents, and generative AI
Why it matters: API habits and workplace usage for Claude, which shows up in both the model layer and day-to-day building.
Why it matters: Model Context Protocol is how an assistant reaches files, APIs, and data without a one-off integration each time.
Why it matters: short, specific labs — prompt engineering, LangChain, fine-tuning, embeddings on Vertex AI, diffusion, and Semantic Kernel — instead of one endless survey.
Why it matters: multi-agent roles and tool use, which is the shape of the predictive maintenance system.
Stack: Python, AutoGen.
Why it matters: research agents that route questions, summarize, and debug tool calls over your own documents.
Why it matters: a functional picture of how generative models work and where companies try to create value with them.
Stack: AWS, Python.
Foundations
Why it matters: regression, trees, clustering, bias and variance, regularization, neural nets, and transformers are the vocabulary under every later course.
Software side: Python, data structures, TensorFlow or PyTorch, and scikit-learn.
Why it matters: instructor-led labs on training, CNNs, data augmentation, RNNs, autoencoders, and GANs.
Stack: GPU notebook, JupyterLab, TensorFlow, Keras.
Why it matters: Python plus AI-assisted coding, which is how a lot of the small apps on the work page were built.
Why it matters: reviews, CI, testing, and production habits around the models, not only the models themselves.
Why it matters: a single map of paths, tools, and communities when the course list starts to sprawl.
Reading
Systems
Why it matters: Kleppmann’s three pillars — reliability, scalability, maintainability — are the checklist I use when a demo has to survive real data.
Why it matters: a repeatable frame and case studies for systems I have not operated at that scale yet.
Why it matters: Chip Huyen treats data, features, retraining, and monitoring as one system, which matches how infrastructure people already think.
Why it matters: a 7-step approach and worked examples (smart compose, personalized generation) for designing GenAI systems, not only calling an API.
Why it matters: ten end-to-end ML system questions with diagrams, useful when a project needs a boundary and a metric, not another notebook.
Strategy and context
Why it matters: a long view of information networks, from bureaucracy to AI, which is the backdrop for the industry stack on the homepage.
Why it matters: a first-person account of how modern computer vision, and a lot of today’s AI, actually got built.
Why it matters: the friction between business goals and technical goals, which is also the subject of the doctoral research interest.
Why it matters: Nigel Vaz’s SPEED frame — strategy, product, engineering, experience, data — is a useful checklist when AI work has to land in an existing company.
Why it matters: product language for shipping models, which complements the infrastructure background rather than replacing it.
Papers and articles
Signals
Consulting-firm posts and arXiv-style papers still load from the existing APIs. They are reference shelves, not the front door.
Latest posts from McKinsey, BCG, Bain, and other firms. Useful as a scan, thin without a personal note on each item.
Open the feedA paper feed across AI, vision, language, and related tags. Same idea: a scan, not a curated essay.
Open the feed