Extending Direct Nash Optimization For Regularized Preferences

Extending Direct Nash Optimization for Regularized Preferences | HackerNoon

Last updated: 2025/04/17 at 1:56 PM

News Room Published 17 April 2025

Authors:

(1) Corby Rosset, Microsoft Research and Correspondence to [email protected];

(2) Ching-An Cheng, Microsoft Research;

(3) Arindam Mitra, Microsoft Research;

(4) Michael Santacroce, Microsoft Research;

(5) Ahmed Awadallah, Microsoft Research and Correspondence to [email protected];

(6) Tengyang Xie, Microsoft Research and Correspondence to [email protected].

Table of Links

Abstract and 1 Introduction

2 Preliminaries

2.1 RLHF Based on Reward Models

2.2 RLHF with General Preferences

3 Direct Nash Optimization and 3.1 Derivation of Algorithm 1

3.2 Theoretical Analysis

4 Practical Algorithm – Iterative Contrastive Self-Improvement

5 Experiments and 5.1 Experimental Setup

5.2 Results and Analysis

6 Related Work

7 Conclusion and References

Appendix

A Extension to Regularized Preferences

B Detailed Proofs

C Additional Experimental Details

A Extension to Regularized Preferences

In this section, we discuss how to extend the DNO framework to the case of regularized preferences (defined in Eq. (5)),

which was first introduced and solved by Munos et al. (2023) via Nash-MD introduced earlier.

This paper is available on arxiv under CC BY 4.0 DEED license.

Extending Direct Nash Optimization for Regularized Preferences | HackerNoon

Table of Links

A Extension to Regularized Preferences

Leave a Reply Cancel reply

Stay Connected

Latest News

Microsoft’s Azure Linux 3.0.20250702 Brings Many Security Fixes

Adaptive Power in iOS 26 Could Mean Longer Stretches Between iPhone Charges

It’s hard to get excited by the Galaxy Watch 8 launch when the Watch 7 is this cheap

ROMOSS suspends production for six months · TechNode

World of Software is your one-stop website for the latest tech news and updates, follow us now to get the news that matters to you.

Quick Link

Topics

Sign Up for Our Newsletter

Table of Links

A Extension to Regularized Preferences

Sign Up For Daily Newsletter

Be keep up! Get the latest breaking news delivered straight to your inbox.

Leave a Reply Cancel reply

Stay Connected

Latest News