Session Overview
From Model Overload to Small, Local, and Multi-Agent AI.
Sumit Sharma opened by sizing up just how fast Generative AI is moving — with over 378,000 language models listed on Hugging Face, and a single 10-day window in which Microsoft, Meta, Databricks, and Google all shipped major new models (Phi-3, LLaMA 3, DBRX, and Gemini 1.5 Pro respectively). That pace, he argued, is exactly why enterprises struggle: in 2023 only about 10% of Generative AI pilots made it to production, largely because organizations lacked "AI-ready data," leading to data drift, high infrastructure costs, and frustrating latency once real enterprise data entered the picture.
He walked through the "build vs. buy" decision enterprises face — SaaS models like ChatGPT are easy to adopt but behave as a black box, which doesn't sit well with AI ethics officers who need to know what data a model was trained on. This sets up the core theme of the talk: the shift toward Small Language Models (SLMs) that run at the edge — directly on factory floor machines, IoT devices, or laptops — which matters enormously for regulated sectors like defense and security that can't send data to the cloud.
Using Microsoft's Phi-3 as a concrete example, he showed how a 3.8-billion-parameter model, focused on data quality over data quantity and shrunk via quantization to fit in just 2GB of RAM, can run locally on a device like an iPhone 14 — while still offering a 4K context window extendable to 128K via "long rope." He closed by describing the road ahead: rather than one giant LLM doing everything, enterprises will orchestrate multiple specialized SLMs (augmented with RAG for missing knowledge) in multi-agent systems — and leaders need to treat Generative AI as a long-term strategic investment, not judge it on one or two short-term use cases, with the end goal being augmented humans freed up for high-value work.
Key Takeaways & Concepts
Session Highlights