Skip to main content

Multimodal AI Engineering

COM SCI 910.5

A six‑week, Python‑driven applied AI engineering course where you build and deploy multimodal systems in vision, voice, and document intelligence, gaining practical skills, evaluations, and a portfolio of six live demos.

Get More Info

 

About This Course

Build and deploy six AI applications in six weeks: a semantic image search engine, a benchmarked visual question‑answering app, a real‑time voice agent, a document Q&A system with cited answers over real PDFs, a cross‑modal generative studio, and a capstone you design. This applied AI engineering course for working professionals centers on Vision, Voice, and Document Intelligence with Python, using the current generation of vision, speech, retrieval‑augmented generation (RAG), and generative AI models. Each week concludes with a fully working system deployed to a public portfolio on Hugging Face Spaces.

The curriculum teaches the decisions practitioners face on the job: frontier APIs versus open‑weight models, quality versus cost and latency, and evaluation the way industry runs it, including LLM‑as‑judge protocols and groundedness testing. Every assignment receives individualized written feedback from the instructor, and weekly demo threads put your deployed work in front of a cohort of peers for structured review. You leave with six live demos, a decision framework for multimodal system design, and the vocabulary to lead these projects at work.