BEGIN:VCALENDAR
PRODID:-//Columba Systems Ltd//NONSGML CPNG/SpringViewer/ICal Output/3.3-
 M3//EN
VERSION:2.0
CALSCALE:GREGORIAN
METHOD:PUBLISH
BEGIN:VEVENT
DTSTAMP:20250915T162351Z
DTSTART:20251112T110000Z
DTEND:20251112T120000Z
SUMMARY:AI-Fun & ELLIS Invited Speaker Series | Chengchun Shi
UID:{http://www.columbasystems.com/customers/uom/gpp/eventid/}cot-mffnqi7
 a-qg6o76
DESCRIPTION:November’s AI-Fun and ELLIS Invited Speaker lecture will take
  place in Nancy Rothwell Building room 3A.012. This is in Core 1\, on th
 e third floor. If you go by stairs\, once you get to the third floor sta
 irwell/seating area\, .012 is the first room on the right (entrance down
  the corridor on the right). If you use the left\, head left out of the 
 lift and room .012 is ahead of you on the left (entrance down the corrid
 or on the right).\n\nNovember’s speaker is Chengchun Shi from the London
  School of Economics.\n\nBio: Chengchun is an associate professor of dat
 a science at London School of Economics and Political Science (LSE). Che
 ngchun completed his PhD in Statistics at North Carolina State Universit
 y before moving to LSE as assistant professor of data science. Chengchun
  has received the Peter Gavin Hall Institute of Mathematical Statistics 
 (IMS) Early Career Prize\, IMS Tweedie Award and the Royal Statistical S
 ociety (RSS) Research Prize.\n\nTalk Title: Doubly Robust Alignment for 
 Large Language Models\n\nAbstract: This talk focuses on reinforcement le
 arning from human feedback (RLHF) for aligning large language models wit
 h human preferences. While RLHF has demonstrated promising results\, man
 y algorithms are highly sensitive to misspecifications in the underlying
  preference model (e.g.\, the Bradley-Terry model)\, the reference polic
 y\, or the reward function\, resulting in undesirable fine-tuning. To ad
 dress model misspecification\, we propose a doubly robust preference opt
 imization algorithm that remains consistent when either the preference m
 odel or the reference policy is correctly specified (without requiring b
 oth). Our proposal demonstrates superior and more robust performance tha
 n state-of-the-art algorithms\, both in theory and in practice. The code
  is available at https: //github.com/DRPO4LLM/DRPO4LLM \n\nIn-person att
 endance is encouraged but if you can only join online\, you are welcome 
 to register via Ticket Source (link provided on this page).
STATUS:TENTATIVE
TRANSP:TRANSPARENT
CLASS:PUBLIC
LOCATION:3A.012\, Nancy Rothwell Building\, Booth Street East\, Mancheste
 r\, M13 9PL
END:VEVENT
END:VCALENDAR
