Optimisation ci politik yu jege
Politigu Optimisation bu Jege (PPO) mooy algorithm biy gëna dooleel jàngat bi gëna méngoo ak modeli làkk yu gëna jubal ci feedback nit.
Résumé
It improves a policy in careful, small steps to avoid the instability that plagues naive policy gradient methods.
Plongeur bu xóot
PPO was introduced by OpenAI in 2017 and became the workhorse behind RLHF for systems like InstructGPT and ChatGPT. The core challenge in policy-gradient RL is that a single overly large update can collapse performance. PPO addresses this with a 'clipped surrogate objective': it measures how much more (or less) likely an action has become versus the old policy, multiplies that ratio by the advantage (how much better the action was than expected), and clips the ratio to a small range like 0.8 to 1.2. This caps how far the policy can move per update, keeping learning stable while still allowing steady improvement. In language-model RLHF, the 'action' is generating a token or response, the reward comes from a reward model, and a KL-divergence penalty keeps the model from drifting too far from its original behavior.
Gis-gis xarala
PPO dafay gëna yokk mébet buñ dagg: min (ratio * njariñ, dagg (ratio, 1-eps, 1 + eps) * njariñ), fu ratio mooy probabilite jëf ju bees-ci-màgget. Njariñ yi dañu leen di faral di xayma ci xayma njariñ yu mat sëkk ak reso valeur jàngat (critique). Ci RLHF, neexal bi dafay boole poñ yi modelu neexal bi am ak penalti KL bu token bu nekk ci politiku royuwaay bi, di ekilibre benefiis yi ci nekk ci wetu modelu njëkk bi.
njeextalu pexe
Gaawaay ak yaatuwaay
Liggéeyukaay yi ci làkk yi mën nañu gëna gaaw te duñu yàq deggoo gi.
Dugg ak yegg
Dafay yaatal jëfandikoo gi ci làkk yi ak ci anam yi ñuy jokkoo.
dogal yu gëna leer
Ekip yi mën nañu gëna yàgg ci àtte ci jamono ji otomatisation di liggéey ci baamtu.
Ëlëgu gëna xéewale ci wàllu politik
PPO mingi wéy di am doole waaye dafa xawa jafe: fàww mu am reso bu wuute, tuning hyperparamètre bu baax, ak ordinatër yu bari. Alternatif yu gëna yomba ñu ngi gëna am doole, lu ci melni DPO (amul RL) ak GRPO, biy wàññi reso valeur bi ci xayma njariñ yi ci kuréel yu tontu yuñ jël misaal, ba noppi dooleel xeetu xalaat yu bees yi. PPO dina wéy fépp fu gëstu ci politik di jàppale dëgg, waaye barab bi mingi jënd ak jaay yenn ci jafe-jafe yi ci pexe yu gëna xéewale.
Doxal ci àdduna dëgg
Reglage InstructGPT ak ChatGPT ngir topp tegtal yi ak tànneefi nit ñi jaaraleko ci RLHF
Taggat ndawu jeu-jeu ak robotik, domen bu njëkk bu PPO balaa xeetu làkk yi
Wàññi toxisite wala gëna mëna jàppale ci yokk poñ yi ci modelu neexal ci suufu KL constraint
Optimiser jëfandikoo jumtukaay wala doxalinu ndawu liggéey bu bari jéego fu ñuy neexal ab model ndax def ay liggéey ci anam wu jaar yoon
Risk yi ak balustrade yi
Lépp lu jaarul yoon mën na dugg ci rapoor yi, jàppale ci liggéey bi, wala ci njariñu gëstu bi.
Sensibilite bu gaaw mën na jur njariñ yu wuute ci laajte yu noonu mel.
Done yu am solo mën nañu feeñ sudee seytu jëfandikoo gi néew doole.
Roadmap ngir samp gi
Mandargal formaa génne gi, melokaan bi, ak standard kalite yi laata ngay dugal ko.
Tontu yu am solo ak balluwaay yu wóor saa yu dëggu bi di am solo.
Fexeel am barabu xool nit ñi ngir am njariñ yu am solo.
Toppal anami gacce yi ak di faral di tàggataat ay laaj wala def-liggéey.
Weyal di banneexu
Free newsletter
Get the daily AI briefing
Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.
One email each weekday. Unsubscribe in one click. We never sell or share your address.
Test yourself
Take the Proximal Policy Optimization quiz
Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.
Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation
Gis bi ci topp
Groupe bi gëna xéewale ci wàllu politik
Laaj yi ñuy faral di laaj
What is Proximal Policy Optimization?
Politigu Optimisation bu Jege (PPO) mooy algorithm biy gëna dooleel jàngat bi gëna méngoo ak modeli làkk yu gëna jubal ci feedback nit. Dafay gëna suqali politik ci jéego yu ndaw, yu moytu ngir moytu ñàkka dal giy sonal pexe gradient politik yu naïf yi.
Ban jafe-jafe la 'dagg' PPO di njëkka saafara?
Dagg ratio probabilite bi dafay tere benn yeesal toxal politik bi fu sori, te loolu mooy gëna indi ñàkka dal ci pexe politik-gradient.
Ci xeetu làkk RLHF ak PPO, fan la siñaalu neexal bi di faral di jogee?
Benn xeetu neexal dafay joxe neexal scalar ngir tontu yiñ defar, ndax nit ñi dañuy poñ bépp génne ci RL du mënati am.
Lan la penalti KL-divergence di def ci RLHF bu sukkandiko ci PPO?
Penalty KL bu nekk ci kaw politiku (royuwaay) bi njëkk, dafay tere model bi soppi doxalinam bu baax bi muy topp neexal bi.
Lan mooy wareefu reso valeur (critique) ci PPO?
PPO xeetu aktër-kritik la; Reseau valeur bi dafay xayma neexal bi ñuy seentu, loolu mooy tax ñu mëna xayma njariñ (dafay faral di jaar ci GAE) luy wàññi variance.
Luy 'njariñ' ci PPO?
Njariñ li mooy natt ba ñaata la jëf jiñ tànn gëna baax (wala gëna bon) buñu ko méngale ak xayma valeur bi, di wax politik bi jëf yi mu wara dooleel.