
Beyond Accuracy: Evaluating the learned representations of Generative AI models | Aida Nematzadeh
Jay Shah16 June 2026Watch on YouTube
Description
Dr. Aida Nematzadeh is a Senior Staff Research Scientist at Google DeepMind where her research focused on multimodal AI models. She works on developing evaluation methods and analyze model’s learning abilities to detect failure modes and guide improvements. Before joining DeepMind, she was a postdoctoral researcher at UC Berkeley and completed her PhD and Masters in Computer Science from the University of Toronto. During her graduate studies she studied how children learn semantic information through computational (cognitive) modeling. Time stamps of the conversation 00:00 Highlights 01:20 Introduction 02:08 Entry point in AI 03:04 Background in Cognitive Science & Computer Science 04:55 Research at Google DeepMind 05:47 Importance of language-vision in AI 10:36 Impact of architecture vs. data on performance 13:06 Transformer architecture 14:30 Evaluating AI models 19:02 Can LLMs understand numerical concepts 24:40 Theory-of-mind in AI 27:58 Do LLMs learn theory of mind? 29:25 LLMs as judge 35:56 Publish vs. perish culture in AI research 40:00 Working at Google DeepMind 42:50 Doing a Ph.D. vs not in AI (at least in 2025) 48:20 Looking back on research career More about Aida: http://www.aidanematzadeh.me/ About the Host: Jay is a Machine Learning Engineer at PathAI working on improving AI for medical diagnosis and prognosis. Linkedin: https://www.linkedin.com/in/shahjay22/ Twitter: https://twitter.com/jaygshah22 Homepage: https://jaygshah.github.io/ for any queries. Stay tuned for upcoming webinars! ***Disclaimer: The information in this video represents the views and opinions of the speaker and does not necessarily represent the views or opinions of any institution. It does not constitute an endorsement by any Institution or its affiliates of such video content.***
What you'll learn
- Evaluation of generative AI models extends beyond accuracy alone; understanding learned representations and failure modes is critical.
- Multimodal AI combining language and vision requires thorough investigation of how models process semantic information across modalities.
- The distinction between architecture and data impact on model performance is crucial for understanding what models truly learn and how to improve them.
- Theory-of-mind in AI examines how models understand and predict concepts like intentionality and mental states.
- Cognitive research on how children learn semantic information offers insights for evaluating AI model capabilities.
Frequently asked questions
What are the key evaluation methods for generative AI models according to Dr. Nematzadeh?
Why is multimodal AI (language-vision combination) important in AI research?
Can LLMs truly understand numerical concepts according to Dr. Nematzadeh's research?
What is the importance of cognitive modeling for AI research?
Topics
In this video
Related reads
OpenAI mikt op lang draaiende agents met overname van Ona
OpenAI neemt Ona over, voorheen bekend als Gitpod, een startup die AI-agents laat draaien in cloud sandboxes. De overname versterkt Codex en stelt OpenAI in staat agents taken te laten uitvoeren die uren of dagen in beslag nemen.
OpenAI introduceert ChatGPT Work, een AI-agent die taken zelfstandig afrondt
OpenAI heeft ChatGPT Work gelanceerd, een AI-agent binnen ChatGPT die complexe projecten zelfstandig uitvoert en eindproducten oplevert zoals spreadsheets, presentaties en webapplicaties. De lancering gaat gepaard met het samenvoegen van de Codex-app in de ChatGPT-desktopapp en de introductie van de GPT-5.6-modelreeks.
OpenAI lanceert ChatGPT Work, agent voor workflows in Google Drive en Slack
OpenAI presenteert ChatGPT Work, een agentgebaseerd product met GPT-5.6 dat zelfstandig complexe workflows over meerdere applicaties kan afhandelen.
Teammaite haalt Twinning Participaties binnen als strategische partner
Het Enschedese AI-bedrijf Teammaite heeft een samenwerking gesloten met participatiemaatschappij Twinning Participaties. Met deze stap wil het in 2025 opgerichte bedrijf de doorontwikkeling van zijn AI-agenten voor softwareteams versnellen.
OpenAI wil api-prijzen verlagen om concurrentie met Anthropic bij te houden
OpenAI overweegt de tarieven voor zijn api flink te verlagen, zo melden bronnen aan The Wall Street Journal. Het bedrijf reageert daarmee op verwachte prijsverlagingen van concurrent Anthropic.
Bijna vier op de tien Europeanen gebruikt AI bij het oriënteren op aankopen
38 procent van de Europeanen zet generatieve AI in om producten te onderzoeken voordat ze een aankoopbeslissing nemen. Dat blijkt uit onderzoek waarover Emerce bericht.