1. EachPod

Deep Papers - Podcast

Deep Papers

Deep Papers is a podcast series featuring deep dives on today’s most important AI papers and research. Hosted by Arize AI founders and engineers, each episode profiles the people and techniques behind cutting-edge breakthroughs in machine learning. 

Technology Science Mathematics Business
Update frequency
every 15 days
Average duration
36 minutes
Episodes
55
Years Active
2023 - 2025
Share to:
Small Language Models are the Future of Agentic AI

Small Language Models are the Future of Agentic AI

We had the privilege of hosting Peter Belcak – an AI Researcher working on the reliability and efficiency of agentic systems at NVIDIA – who walked us through his new paper making the rounds in AI ci…

00:31:15  |   Fri 05 Sep 2025
Watermarking for LLMs and Image Models

Watermarking for LLMs and Image Models

In this AI research paper reading, we dive into "A Watermark for Large Language Models" with the paper's author John Kirchenbauer. 

This paper is a timely exploration of techniques for embedding invis…

00:42:56  |   Wed 30 Jul 2025
Self-Adapting Language Models: Paper Authors Discuss Implications

Self-Adapting Language Models: Paper Authors Discuss Implications

The authors of the new paper *Self-Adapting Language Models (SEAL)* shared a behind-the-scenes look at their work, motivations, results, and future directions.

The paper introduces a novel method for …

00:31:26  |   Tue 08 Jul 2025
The Illusion of Thinking: What the Apple AI Paper Says About LLM Reasoning

The Illusion of Thinking: What the Apple AI Paper Says About LLM Reasoning

This week we discuss The Illusion of Thinking, a new paper from researchers at Apple that challenges today’s evaluation methods and introduces a new benchmark: synthetic puzzles with controllable com…

00:30:35  |   Fri 20 Jun 2025
Accurate KV Cache Quantization with Outlier Tokens Tracing

Accurate KV Cache Quantization with Outlier Tokens Tracing

We discuss Accurate KV Cache Quantization with Outlier Tokens Tracing, a deep dive into improving the efficiency of LLM inference. The authors enhance KV Cache quantization, a technique for reducing …

00:25:11  |   Wed 04 Jun 2025
Scalable Chain of Thoughts via Elastic Reasoning

Scalable Chain of Thoughts via Elastic Reasoning

In this week's episode, we talk about Elastic Reasoning, a novel framework designed to enhance the efficiency and scalability of large reasoning models by explicitly separating the reasoning process …

00:28:54  |   Fri 16 May 2025
Sleep-time Compute: Beyond Inference Scaling at Test-time

Sleep-time Compute: Beyond Inference Scaling at Test-time

What if your LLM could think ahead—preparing answers before questions are even asked?

In this week's paper read, we dive into a groundbreaking new paper from researchers at Letta, introducing sleep-ti…

00:30:24  |   Fri 02 May 2025
LibreEval: The Largest Open Source Benchmark for RAG Hallucination Detection

LibreEval: The Largest Open Source Benchmark for RAG Hallucination Detection

For this week's paper read, we dive into our own research.

We wanted to create a replicable, evolving dataset that can keep pace with model training so that you always know you're testing with data yo…

00:27:19  |   Fri 18 Apr 2025
AI Benchmark Deep Dive: Gemini 2.5 and Humanity's Last Exam

AI Benchmark Deep Dive: Gemini 2.5 and Humanity's Last Exam

This week we talk about modern AI benchmarks, taking a close look at Google's recent Gemini 2.5 release and its performance on key evaluations, notably  Humanity's Last Exam (HLE). In the session we …

00:26:11  |   Fri 04 Apr 2025
Model Context Protocol (MCP)

Model Context Protocol (MCP)

We cover Anthropic’s groundbreaking Model Context Protocol (MCP). Though it was released in November 2024, we've been seeing a lot of hype around it lately, and thought it was well worth digging into…

00:15:03  |   Tue 25 Mar 2025
AI Roundup: DeepSeek’s Big Moves, Claude 3.7, and the Latest Breakthroughs

AI Roundup: DeepSeek’s Big Moves, Claude 3.7, and the Latest Breakthroughs

This week, we're mixing things up a little bit. Instead of diving deep into a single research paper, we cover the biggest AI developments from the past few weeks.

We break down key announcements, incl…

00:30:23  |   Sat 01 Mar 2025
How DeepSeek is Pushing the Boundaries of AI Development

How DeepSeek is Pushing the Boundaries of AI Development

This week, we dive into DeepSeek. SallyAnn DeLucia, Product Manager at Arize, and Nick Luzio, a Solutions Engineer, break down key insights on a model that have dominating headlines for its significa…

00:29:54  |   Fri 21 Feb 2025
Multiagent Finetuning: A Conversation with Researcher Yilun Du

Multiagent Finetuning: A Conversation with Researcher Yilun Du

We talk to Google DeepMind Senior Research Scientist (and incoming Assistant Professor at Harvard), Yilun Du, about his latest paper, "Multiagent Finetuning: Self Improvement with Diverse Reasoning C…

00:30:03  |   Tue 04 Feb 2025
Training Large Language Models to Reason in Continuous Latent Space

Training Large Language Models to Reason in Continuous Latent Space

LLMs have typically been restricted to reason in the "language space," where chain-of-thought (CoT) is used to solve complex reasoning problems. But a new paper argues that language space may not alw…

00:24:58  |   Tue 14 Jan 2025
LLMs as Judges: A Comprehensive Survey on LLM-Based Evaluation Methods

LLMs as Judges: A Comprehensive Survey on LLM-Based Evaluation Methods

We discuss a major survey of work and research on LLM-as-Judge from the last few years. "LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods" systematically examines the LLMs-as-Ju…

00:28:57  |   Mon 23 Dec 2024
Merge, Ensemble, and Cooperate! A Survey on Collaborative LLM Strategies

Merge, Ensemble, and Cooperate! A Survey on Collaborative LLM Strategies

LLMs have revolutionized natural language processing, showcasing remarkable versatility and capabilities. But individual LLMs often exhibit distinct strengths and weaknesses, influenced by difference…

00:28:47  |   Tue 10 Dec 2024
Agent-as-a-Judge: Evaluate Agents with Agents

Agent-as-a-Judge: Evaluate Agents with Agents

This week, we break down the “Agent-as-a-Judge” framework—a new agent evaluation paradigm that’s kind of like getting robots to grade each other’s homework. Where typical evaluation methods focus sol…

00:24:54  |   Sat 23 Nov 2024
Introduction to OpenAI's Realtime API

Introduction to OpenAI's Realtime API

We break down OpenAI’s realtime API. Learn how to seamlessly integrate powerful language models into your applications for instant, context-aware responses that drive user engagement. Whether you’re …

00:29:56  |   Tue 12 Nov 2024
Swarm: OpenAI's Experimental Approach to Multi-Agent Systems

Swarm: OpenAI's Experimental Approach to Multi-Agent Systems

As multi-agent systems grow in importance for fields ranging from customer support to autonomous decision-making, OpenAI has introduced Swarm, an experimental framework that simplifies the process of…

00:46:46  |   Tue 29 Oct 2024
Disclaimer: The podcast and artwork embedded on this page are the property of Arize AI. This content is not affiliated with or endorsed by eachpod.com.