<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Papers on Vishal Bhoriwal</title><link>https://www.vishalbhoriwal.me/papers/</link><description>Recent content in Papers on Vishal Bhoriwal</description><generator>Hugo</generator><language>en</language><lastBuildDate>Wed, 22 Jan 2025 00:00:00 +0000</lastBuildDate><atom:link href="https://www.vishalbhoriwal.me/papers/index.xml" rel="self" type="application/rss+xml"/><item><title>DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning</title><link>https://www.vishalbhoriwal.me/papers/deepseek-r1/</link><pubDate>Wed, 22 Jan 2025 00:00:00 +0000</pubDate><guid>https://www.vishalbhoriwal.me/papers/deepseek-r1/</guid><description>&lt;p&gt;&lt;strong&gt;Link:&lt;/strong&gt; &lt;a href="https://arxiv.org/abs/2501.12948" target="_blank" rel="noopener noreferrer"&gt;arxiv.org/abs/2501.12948&lt;/a&gt;
&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="abstract"&gt;Abstract&lt;/h2&gt;
&lt;p&gt;DeepSeek-R1 shows that strong reasoning capabilities in large language models can emerge from pure reinforcement learning — no human-labeled reasoning traces required. The resulting model matches OpenAI&amp;rsquo;s o1 on a range of benchmarks, and the learned reasoning patterns can be distilled into much smaller models that still punch well above their weight.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id="background"&gt;Background&lt;/h2&gt;
&lt;p&gt;Getting LLMs to reason reliably has mostly meant one thing: more human supervision. Chain-of-thought prompting helps, but it requires annotated examples. RLHF aligns models to human preferences, but again depends on human feedback. OpenAI&amp;rsquo;s o1 demonstrated that models could &amp;ldquo;think&amp;rdquo; before answering using extended reasoning chains — but the training details were never released.&lt;/p&gt;</description></item></channel></rss>