Image Processing and Computer Vision: From Pixels to Intelligent Visual Recognition

November
10
2026 (Tuesday)
Time 08:00 AM PST | 11:00 AM EST
Duration: 60 Minutes
49 Days Left To REGISTER
Id: 213516
Instructor
Mohammed Rizwan Roshan 
Looking for customized training for your team? Please contact our Customer Support Team (support@masterlearning.com) to discuss your requirements.
Live
Recorded
Live + Recorded

Overview

This webinar provides an accessible introduction to Image Processing and Computer Vision, beginning with the representation of images and progressing toward object recognition and modern deep-learning-based vision systems.

The session first introduces spatial filtering, including Gaussian and median filtering, and explains how filtering can be used for smoothing, noise suppression, and preservation of important image structures.

The webinar then introduces edge detection, explaining gradients, gradient magnitude and direction, and the importance of non-maximum suppression in obtaining thin and accurately localized edges.

The session subsequently introduces the image formation model and camera calibration, including intrinsic and extrinsic parameters, coordinate systems, projection, lens distortion, and Zhang's camera calibration approach.

The next stage focuses on local feature detection and description. Scale-space and Difference of Gaussian are used to understand how keypoints can be detected at characteristic scales. SIFT is then introduced as a scale- and rotation-invariant feature descriptor. The webinar explains canonical orientation, local coordinate frames, orientation histograms, descriptor construction, and feature matching.

The discussion then moves to instance-level object detection, covering template matching, shape-based matching, Hough Transform, Generalized Hough Transform, and model-based feature matching. These methods demonstrate how a known object can be located in a new image even when its position, scale, or orientation changes.

Finally, the webinar connects classical Computer Vision to Deep Learning and CNN architectures, introducing the evolution from LeNet to AlexNet, ZFNet, VGG, GoogleNet/Inception, and ResNet. Important architectural ideas such as pooling, convolutional layers, grouped convolutions, bottleneck blocks, skip connections, global average pooling, and computational efficiency are discussed.

The overall goal is to provide participants with a coherent understanding of how a computer progresses from raw pixels to meaningful visual information and object recognition.

Why you should Attend

AI systems can recognize faces, objects, scenes, and visual patterns-but what actually happens between an image entering the computer and an object being recognized?

Understanding Computer Vision requires more than knowing that "CNNs recognize images." The underlying ideas begin much earlier: filtering removes unwanted information, gradients reveal changes, edges identify boundaries, local features provide distinctive points, descriptors represent local appearance, and matching establishes correspondences between images.

This webinar will help participants understand the complete evolution of visual recognition, from classical image-processing techniques to modern deep-learning architectures.

Participants will also understand why techniques such as SIFT, Hough Transform, and camera calibration remain important for understanding the principles behind modern computer vision.

Areas Covered in the Session

  • Introduction to images and image processing 
  • Spatial filtering and noise reduction 
  • Edge detection and Non-Maximum Suppression 
  • Image formation and camera calibration 
  • Scale-space and Difference of Gaussian (DoG) 
  • SIFT keypoints and feature descriptors 
  • Feature matching and Lowe's ratio test 
  • Template Matching and Hough Transform 
  • Generalized Hough Transform and instance detection 
  • CNNs: LeNet, AlexNet, VGG, Inception and ResNet 
  • Convolution, pooling, Inception modules and residual connections 
  • Evolution from classical Computer Vision to Deep Learning

Who Will Benefit

  • Computer Science Students
  • Engineering Students
  • Artificial Intelligence / Machine Learning Students
  • Computer Vision Students
  • Robotics Students
  • Data Science Students
  • Software Developers interested in AI and Computer Vision
  • Aspiring AI/ML Engineers
  • Students studying Image Processing
  • Students studying Pattern Recognition
  • Researchers interested in Computer Vision
  • Students preparing for Computer Vision and AI examinations
  • Anyone interested in how machines understand images

Speaker Profile

Mohammed Rizwan Roshan is a Computer Science graduate with strong hands-on experience in software development, mobile application development, and Machine Learning. He has worked at Zoho Corporation, contributing to SaaS-based systems and gaining exposure to production-level software development. Beyond enterprise software, he has extensive experience building end-to-end applications, ranging from small-scale prototypes to fully deployed, user-facing production systems. This includes developing cross-platform mobile and web applications, several of which are actively used by organizations and users. He has also worked on multiple Machine Learning projects, applying Python-based ML techniques to real datasets. This practical ML experience is complemented by academic training, as he is currently pursuing a Masters degree in Artificial Intelligence, with exposure to core ML concepts, neural networks, NLP, and data-driven problem solving.

In addition, Rizwan Roshanhas experience in Cybersecurity fundamentals, and has presented technical papers on Google Firebase and Mobile Application Development at academic events. Having led development teams and participated in national-level competitions, He brings a balanced perspective that connects Computer Science fundamentals, Machine Learning concepts, real-world implementation, and career relevance - making complex AI topics accessible, practical, and industry-oriented.
Access Recorded Version
One Attendee / Group Attendees

Unlimited Viewing Recorded Version for 6 months ( Access information will be emailed 24 hours after the completion of live webinar)