This two-day summer school provides participants with a practical understanding of how modern artificial intelligence workloads are developed, scaled, deployed, and optimized on high-performance computing systems.
The program combines HPC architecture, Linux and Slurm-based job management, computer vision, large language models, multimodal AI, distributed training, inference optimization, and AI compiler technologies. Through expert-led lectures, demonstrations, and guided practical sessions, participants will gain exposure to the complete AI engineering workflow—from preparing and training models to deploying them efficiently on GPU-enabled supercomputing platforms.
Learning Outcomes
Upon completion of the summer school, participants will be able to:
Explain the architecture and operation of GPU-enabled HPC systems.
Work in a Linux-based HPC environment and manage jobs using Slurm.
Prepare and execute computer-vision, language, and multimodal AI workloads.
Explain single-GPU, multi-GPU, and multi-node training approaches.
Understand distributed-training strategies and communication overheads.
Develop introductory LLM workflows using parameter-efficient fine-tuning and retrieval-augmented generation.
Evaluate inference performance in terms of latency, throughput, memory utilization, and batching.
Explain how MLIR, LLVM, and runtime systems map AI models onto computing hardware.
Audience
Who Should Attend
The summer school is suitable for:
Undergraduate and graduate students
Researchers and academic faculty
AI and machine-learning engineers
Computer-vision and NLP practitioners
Software and HPC professionals
Engineers working in distributed AI, deployment, or compilers
Entry Requirements
Prerequisites
Participants should have:
Working knowledge of Python
Basic understanding of machine learning or deep learning
Familiarity with the Linux command line
Experience with HPC, Slurm, distributed training, MLIR, or LLVM is useful but not required.
Laptop Requirement Bring a laptop with at least 8 GB RAM and administrator access. Setup instructions will be shared before the summer school.
Software environments, data storage, and resource management
Submitting and managing AI workloads using Slurm
PracticalLinux, software environments, and Slurm job management
Foundations of deep learning for computer vision
Classification, object detection, and image segmentation
Data preparation and preprocessing for large datasets
GPU-based training and evaluation
Performance monitoring and resource utilization
PracticalGPU-based vision-model training and evaluation
Transformer architectures and foundations of large language models
Vision-language and multimodal AI systems
Fine-tuning and parameter-efficient methods
LoRA and adapter-based training
Retrieval-augmented generation
Agentic AI and tool-using workflows
Evaluation and real-world applications
DemoLoRA-based fine-tuning and retrieval-augmented generation
ReviewDay 1 review and closing
Day 2 – 30 August 2026
Principles of parallel and distributed training
Data, model, tensor, and pipeline parallelism
Multi-GPU and multi-node training
Distributed data-parallel execution
Communication overhead and synchronization
Scaling efficiency, profiling, and performance bottlenecks
PracticalMulti-GPU training and performance monitoring
AI model-serving architectures
Inference engines and runtime systems
Quantization, pruning, and model compression
Throughput, latency, batching, and memory optimization
GPU resource utilization
Deployment on cluster and edge-computing targets
PracticalQuantization, batching, and inference benchmarking
Identifying high-value AI use cases
AI applications for local industrial and research challenges
Case studies in vision, language, and multimodal AI
Data, infrastructure, and deployment requirements
Moving from proof of concept to operational deployment
Performance, reliability, and maintainability considerations
Role of compilers in the AI software stack
Introduction to intermediate representations
MLIR and LLVM fundamentals for AI workloads
Graph-level and operator-level optimization
Dialects, lowering, and code generation
Runtime systems for AI execution
Hardware mapping and accelerator support
Demonstration of an AI compilation pipeline
DemoMLIR-based model lowering and runtime execution
ClosingCertificate distribution and closing ceremony
Lead Trainer
Dr. Tassadaq Hussain Cheema
Director, Centre for AI & Big Data (CAID)
Expertise Computer architecture, processor-based systems, parallel computing, and hardware–software co-design.
A computer architect and technology leader with more than 20 years of national and international experience spanning research, industrial systems, project leadership, and professional training.
Registration Process
Submit the form Complete the online registration form with the required details.
Verification Places are assigned first-come, first-served after form and payment verification.
Confirmation Successful applicants are notified through email or WhatsApp.