
Beyond Accuracy: Evaluating the learned representations of Generative AI models | Aida Nematzadeh
Jay Shah16 June 2026Watch on YouTube
Description
Dr. Aida Nematzadeh is a Senior Staff Research Scientist at Google DeepMind where her research focused on multimodal AI models. She works on developing evaluation methods and analyze model’s learning abilities to detect failure modes and guide improvements. Before joining DeepMind, she was a postdoctoral researcher at UC Berkeley and completed her PhD and Masters in Computer Science from the University of Toronto. During her graduate studies she studied how children learn semantic information through computational (cognitive) modeling. Time stamps of the conversation 00:00 Highlights 01:20 Introduction 02:08 Entry point in AI 03:04 Background in Cognitive Science & Computer Science 04:55 Research at Google DeepMind 05:47 Importance of language-vision in AI 10:36 Impact of architecture vs. data on performance 13:06 Transformer architecture 14:30 Evaluating AI models 19:02 Can LLMs understand numerical concepts 24:40 Theory-of-mind in AI 27:58 Do LLMs learn theory of mind? 29:25 LLMs as judge 35:56 Publish vs. perish culture in AI research 40:00 Working at Google DeepMind 42:50 Doing a Ph.D. vs not in AI (at least in 2025) 48:20 Looking back on research career More about Aida: http://www.aidanematzadeh.me/ About the Host: Jay is a Machine Learning Engineer at PathAI working on improving AI for medical diagnosis and prognosis. Linkedin: https://www.linkedin.com/in/shahjay22/ Twitter: https://twitter.com/jaygshah22 Homepage: https://jaygshah.github.io/ for any queries. Stay tuned for upcoming webinars! ***Disclaimer: The information in this video represents the views and opinions of the speaker and does not necessarily represent the views or opinions of any institution. It does not constitute an endorsement by any Institution or its affiliates of such video content.***
What you'll learn
- Evaluation of generative AI models extends beyond accuracy alone; understanding learned representations and failure modes is critical.
- Multimodal AI combining language and vision requires thorough investigation of how models process semantic information across modalities.
- The distinction between architecture and data impact on model performance is crucial for understanding what models truly learn and how to improve them.
- Theory-of-mind in AI examines how models understand and predict concepts like intentionality and mental states.
- Cognitive research on how children learn semantic information offers insights for evaluating AI model capabilities.
Frequently asked questions
What are the key evaluation methods for generative AI models according to Dr. Nematzadeh?
Why is multimodal AI (language-vision combination) important in AI research?
Can LLMs truly understand numerical concepts according to Dr. Nematzadeh's research?
What is the importance of cognitive modeling for AI research?
Topics
In this video
Related reads
OpenAI mikt op lang draaiende agents met overname van Ona
OpenAI neemt Ona over, voorheen bekend als Gitpod, een startup die AI-agents laat draaien in cloud sandboxes. De overname versterkt Codex en stelt OpenAI in staat agents taken te laten uitvoeren die uren of dagen in beslag nemen.
Google klaagt Chinese cybercrimegroep aan voor AI-fraude
Google stelt dat een Chinese organisatie AI gebruikt voor frauduleuze activiteiten en stapt daarin samen met de FBI en Amerikaanse providers naar de rechter.
Apple zou in overleg zijn met Anthropic en Google over Siri-alternatieven
Apple overlegt naar verluidt met Anthropic en Google over het beschikbaar stellen van hun AI-chatbots als alternatief voor Siri via een extensiesysteem.
Okta en Google Cloud beveiligen AI-agents
Okta en Google Cloud breiden samenwerking uit met nieuwe veiligheidsoplossingen voor AI-agents en verbeterde browserbescherming.
Europees internetregister RIPE NCC stapt voor 5 miljoen euro af van Amerikaanse clouddiensten
RIPE NCC, het Amsterdamse internetregister voor Europa, het Midden-Oosten en Centraal-Azië, wil zijn technische infrastructuur tussen 2026 en 2028 volledig vernieuwen en daarbij afstappen van AWS, Google en Cloudflare. De geschatte kosten bedragen 5 miljoen euro.
Rappit lanceert agentic AI-platform voor enterprise softwareontwikkeling
Het Nederlandse Rappit heeft een agentic AI-platform voor enterprise applicatieontwikkeling gelanceerd dat AI-agents combineert met menselijk toezicht en governance. Het bedrijf, voortgekomen uit Vanenburg Software, wil daarmee de spanning wegnemen tussen snelheid en controle bij het bouwen van software.