¶àÂÖAgentѵÁ·¹Õµã£¡Ç廪Ê×´´¿ÉÖ´ÐÐÊý¾Ý±Õ»·£¬£¬£¬£¬£¬£¬£¬£¬¿ªÔ´ÓâÔ½GPT-5
2026-02-27 23:38:01

ÐÂÖÇÔª±¨µÀ

±à¼­£ºLRST

¡¾ÐÂÖÇÔªµ¼¶Á¡¿Ç廪ÍŶÓÌá³öEigenDataϵͳ£¬£¬£¬£¬£¬£¬£¬£¬Í¨¹ý¿ÉÖ´ÐÐÊý¾Ý±Õ»·ÓÅ»¯¶àÂÖAgentѵÁ·£¬£¬£¬£¬£¬£¬£¬£¬ÔÚÕæÊµ³¡¾°ÖÐʹ¿ªÔ´Ä£×ÓÌåÏÖµÖ´ïÓë±ÕԴϵͳÏ൱ˮƽ¡£¡£¡£¡£ ¡£¡£¡£¡£Òªº¦ÔÚÓÚѵÁ·Êý¾ÝµÄÎȹÌÐԺͿÉÑéÖ¤ÐÔ£¬£¬£¬£¬£¬£¬£¬£¬È·±£Ä£×ÓÔÚ½»»¥ÖÐÄÜÒ»Á¬Ñ§Ï°ÓÐÓÃÕ½ÂÔ£¬£¬£¬£¬£¬£¬£¬£¬¶ø·ÇÒÀÀµ²»¿É¿¿µÄ½±ÀøÐźš£¡£¡£¡£ ¡£¡£¡£¡£

ÒÑÍùÒ»Ä꣬£¬£¬£¬£¬£¬£¬£¬AgentµÄ¡¸ÄÜÁ¦¾ºÈü¡¹ÏÕЩ×ßµ½ÁËÒ»¸ö¹Õµã£ºµ¥ÂÖ¹¤¾ßŲÓᢶÌÁ´Â·ÍÆÀíµÄÌáÉý»¹ÔÚ¼ÌÐø£¬£¬£¬£¬£¬£¬£¬£¬µ«Ò»µ©½øÈëÕæÊµ¶àÂÖ½»»¥£¬£¬£¬£¬£¬£¬£¬£¬ÏµÍ³×îÏÈ̻¶³öÍêÈ«²î±ðµÄųÈõÐÔ¡£¡£¡£¡£ ¡£¡£¡£¡£

¹¤³ÌÍŶÓÔ½À´Ô½ÆµÈÔµØÓöµ½Í³Ò»ÎÊÌ⣺ģ×ÓÔÚÀëÏ߯À¹ÀÖÐÌåÏÖÕý³££¬£¬£¬£¬£¬£¬£¬£¬µ«Ò»µ©½øÈëÕæÊµ¶àÂÖ½»»¥£¬£¬£¬£¬£¬£¬£¬£¬ÑµÁ·ÐźžÍ×îÏÈÆµÈÔÊ§Õæ¡£¡£¡£¡£ ¡£¡£¡£¡£

Ò»´ÎÒì³£µÄÓû§ÐÐΪ¡¢Ò»´Î¹¤¾ß¹ì¼£ÅÜÆ«£¬£¬£¬£¬£¬£¬£¬£¬¶¼»á°ÑÕû¶ÎrolloutµÄrewardÖ±½Ó¹éÁ㣬£¬£¬£¬£¬£¬£¬£¬×îÖÕ°ÑÇ¿»¯Ñ§Ï°ÍÆÏò¹ýʧƫÏò¡£¡£¡£¡£ ¡£¡£¡£¡£

Ô½À´Ô½¶àµÄÐźÅÅú×¢AgentѵÁ·ÖУº

¶àÂÖTool-Using AgentµÄÉÏÏÞ£¬£¬£¬£¬£¬£¬£¬£¬Ô½À´Ô½È¡¾öÓÚѵÁ·ÐźÅÊÇ·ñ¿É¹éÒò¡¢¿ÉÑéÖ¤£¬£¬£¬£¬£¬£¬£¬£¬¶ø²»µ«ÊÇÄ£×Ó¹æÄ£¡£¡£¡£¡£ ¡£¡£¡£¡£

ÔÚ¦Ó?-benchµÈÕæÊµTool-Using Agent»ù×¼ÖУ¬£¬£¬£¬£¬£¬£¬£¬Ñо¿ÕßÊӲ쵽£¬£¬£¬£¬£¬£¬£¬£¬¶àÂÖAgentÔÚ½øÈëÇ¿»¯Ñ§Ï°½×¶Îºó£¬£¬£¬£¬£¬£¬£¬£¬ÀÖ³ÉÂʲ¢²»×ÜÊÇËæÑµÁ·Íƽø¶ø¿ÝÔïÌáÉý£¬£¬£¬£¬£¬£¬£¬£¬·´¶ø³£ÅãͬÏÔ×Ų¨¶¯£¬£¬£¬£¬£¬£¬£¬£¬ÕâЩ²¨¶¯²¢·ÇÀ´×ÔÄ£×ÓÄÜÁ¦È±·¦£¬£¬£¬£¬£¬£¬£¬£¬¶ø¸ü¶àÔ´ÓÚ³¤Á´Â·½»»¥ÖÐÓû§ÐÐΪ²»ÎȹÌÓë½±ÀøÎó¹éÒòµÄÒ»Á¬·Å´ó¡£¡£¡£¡£ ¡£¡£¡£¡£

Ò»Ïî×îÐÂÑо¿´Óϵͳ²ãÃæÖØ¹¹Á˶àÂÖAgentµÄѵÁ·Á÷³Ì£ºÎ§ÈÆ¿ÉÖ´ÐÐÊý¾ÝÌìÉú¡¢Óû§Ä£×ÓÎȹ̻¯Óëverifier-based½±ÀøÌá³öÁËÒ»Ì×еÄѵÁ··¶Ê½£¬£¬£¬£¬£¬£¬£¬£¬²¢ÔÚ¦Ó?-benchµÄÈý¸öÕæÊµ¹¤¾ßÓòÉÏÍê³ÉÑéÖ¤¡£¡£¡£¡£ ¡£¡£¡£¡£

ÂÛÎÄÁ´½Ó£ºhttps://arxiv.org/abs/2601.22607

ÔÚ²»ÒýÈë¸ü´óÄ£×Ó¹æÄ£µÄÌõ¼þÏ£¬£¬£¬£¬£¬£¬£¬£¬¿ªÔ´Qwen3ϵÁÐÄ£×ÓÔÚÒªº¦³¡¾°ÖÐʵÏÖÁËÏÔÖøÌáÉý£º

AirlineÖÐ73.0%pass?£¬£¬£¬£¬£¬£¬£¬£¬ÓëGemini 3.0 Pro»ù±¾³Öƽ£¬£¬£¬£¬£¬£¬£¬£¬ÏÔןßÓÚGPT-5£¨62.5%£©

TelecomÖÐ98.3%pass?£¬£¬£¬£¬£¬£¬£¬£¬µÖ´ïÄ¿½ñ¹ûÕæµÄ×î¼ÑЧ¹û£¬£¬£¬£¬£¬£¬£¬£¬Áè¼ÝGemini 3.0 Pro¡¢Claude SonnetÓëGPT-5

ÕâЩЧ¹ûÅú×¢£¬£¬£¬£¬£¬£¬£¬£¬½èÖúϵͳ¼¶ÑµÁ··¶Ê½µÄÓÅ»¯£¬£¬£¬£¬£¬£¬£¬£¬¿ªÔ´Ä£×ÓÔÚÕæÊµ¹¤¾ß½»»¥Ê¹ÃüÉϵĿɿ¿ÐÔÒѾ­±»ÍÆÖÁÓëÖ÷Á÷±ÕԴϵһÇÐÒ»Ìݶӡ£¡£¡£¡£ ¡£¡£¡£¡£

¶àÂÖAgentÄÑѵ

²¢²»ÊÇ¡¸²»»áÓù¤¾ß¡¹

ÈôÊÇֻͣÁôÔÚµ¥ÂÖ¹¤¾ßŲÓòãÃæ£¬£¬£¬£¬£¬£¬£¬£¬AgentµÄÎÊÌâ¿´ÆðÀ´²¢²»Öش󡣡£¡£¡£ ¡£¡£¡£¡£

¸ø¶¨ÊäÈ롢ѡÔñ¹¤¾ß¡¢Ö´ÐÐÒ»´Î¡¢·µ»ØÐ§¹û£¬£¬£¬£¬£¬£¬£¬£¬rewardÒ²¿ÉÒÔÖ±½Ó¶ÔÓ¦µ½ÕâÒ»²½ÊÇ·ñÀֳɡ£¡£¡£¡£ ¡£¡£¡£¡£

µ«Ò»µ©°ÑÊÓ½ÇÀ­µ½ÕæÊµµÄ¶àÂÖ½»»¥ÖУ¬£¬£¬£¬£¬£¬£¬£¬ÇéÐξÍÍêÈ«±äÁË¡£¡£¡£¡£ ¡£¡£¡£¡£

¶Ô»°±»À­³¤Îª³¤Á´Â·µÄtrajectory£¬£¬£¬£¬£¬£¬£¬£¬¹¤¾ßŲÓò»ÔÙÊÇÁæØêÊÂÎñ£¬£¬£¬£¬£¬£¬£¬£¬¶øÊÇÓëÓû§·´Ïì½»Ö¯·ºÆð£»£» £»£»£»£»Óû§×´Ì¬Ò²²»ÔÙÊǾ²Ì¬Ìõ¼þ£¬£¬£¬£¬£¬£¬£¬£¬¶øÊÇÔÚ½»»¥Àú³ÌÖÐһֱ̻¶¡¢ÉõÖÁ±¬·¢Æ¯ÒÆ¡£¡£¡£¡£ ¡£¡£¡£¡£

´Ëʱ£¬£¬£¬£¬£¬£¬£¬£¬Agent ÃæÁÙµÄÒѾ­²»ÊÇ¡¸»á²»»áÓù¤¾ß¡¹£¬£¬£¬£¬£¬£¬£¬£¬¶øÊÇÄÜ·ñÔÚÒ»¸öÒ»Á¬×ª±äµÄϵͳÖмá³Ö¾öÒéÒ»ÖÂÐÔ¡£¡£¡£¡£ ¡£¡£¡£¡£

¶øÔÚÏÖʵѵÁ·ÇéÐÎÖУ¬£¬£¬£¬£¬£¬£¬£¬Ä£×ÓÍùÍùÌåÏÖ³öÏÔ×ŵIJ»ÎȹÌÐÔ£¬£¬£¬£¬£¬£¬£¬£¬Ä£×ÓÈÝÒ×ѧƫ£¬£¬£¬£¬£¬£¬£¬£¬ÉõÖÁ·ºÆðЧ¹ûËæÑµÁ·²¨¶¯¡¢ÄÑÒÔÊÕÁ²µÄÎÊÌâ¡£¡£¡£¡£ ¡£¡£¡£¡£

Ñо¿Ð§¹ûÖ¸Ã÷Ö÷ÒªÔµ¹ÊÔ­Óɼ¯ÖÐÔÚÁ½µã£º

1. ȱ·¦ÕæÕý¡¸¿ÉÓá¹µÄѵÁ·Êý¾Ý

ÕæÕý¿ÉÓÃÓÚ¶àÂÖAgentѵÁ·µÄÊý¾Ý£¬£¬£¬£¬£¬£¬£¬£¬±ØÐèͬʱÁýÕÖ£º

¶àÂÖ¶Ô»°+ ¶à²½¹¤¾ßÖ´ÐÐ + Óû§²àÐÅÏ¢Öð²½Í¸Â¶/¸Ä±äÆ«ºÃ¡£¡£¡£¡£ ¡£¡£¡£¡£

ÎÊÌâÔÚÓÚ£¬£¬£¬£¬£¬£¬£¬£¬ÕâÑùµÄÊý¾ÝÔÚÏÖʵÖÐÏÕЩ²»¿ÉÄÜͨ¹ýÈ˹¤±ê×¢¹æÄ£»£» £»£»£»£»¯»ñµÃ¡£¡£¡£¡£ ¡£¡£¡£¡£¶ø×Ô¶¯ºÏ³ÉµÄÊý¾Ý£¬£¬£¬£¬£¬£¬£¬£¬¿´ËÆ»º½âÁËÊý¾ÝϡȱµÄÎÊÌ⣬£¬£¬£¬£¬£¬£¬£¬È´ÒýÈëÁËеÄÒþ»¼¡£¡£¡£¡£ ¡£¡£¡£¡£

ÔÚ´ó×ÚÑù±¾ÖУ¬£¬£¬£¬£¬£¬£¬£¬¹¤¾ßŲÓù켣ÔÚÎı¾²ãÃæ¡¸¿´ÆðÀ´ºÏÀí¡¹£¬£¬£¬£¬£¬£¬£¬£¬µ«Ö»ÒªÕæÕýÖ´ÐÐÒ»±é£¬£¬£¬£¬£¬£¬£¬£¬¾Í»á´¥·¢²»¿ÉÍê³É״̬£¬£¬£¬£¬£¬£¬£¬£¬trajectory ÔÚÖÐ;ʧ°Ü¡£¡£¡£¡£ ¡£¡£¡£¡£

×îÖÕ£¬£¬£¬£¬£¬£¬£¬£¬Agent ѧµ½µÄ²¢²»ÊÇÎȹ̡¢¿É¸´ÏֵŤ¾ßʹÓÃÄÜÁ¦£¬£¬£¬£¬£¬£¬£¬£¬¶øÊÇÒ»ÖÖÍ£ÁôÔÚ±í²ãµÄÕ½ÂÔģʽ£¨surface-level policy£©£¬£¬£¬£¬£¬£¬£¬£¬¼´Ëü¿´ÆðÀ´ÏñÔÚ×öÊ£¬£¬£¬£¬£¬£¬£¬£¬È´ÎÞ·¨ÔÚÕæÊµÏµÍ³ÖÐÅÜͨ¡£¡£¡£¡£ ¡£¡£¡£¡£

2. Óû§Ä£ÄâµÄ²»ÎȹÌÐÔ»áÖ±½ÓÎÛȾRLÐźÅ

ÔÚinteractive RLÉèÖÃÖУ¬£¬£¬£¬£¬£¬£¬£¬Óû§Ä£ÄâÆ÷ÊÇÇý¶¯¶Ô»°²»¿É»òȱµÄÒ»»·¡£¡£¡£¡£ ¡£¡£¡£¡£µ«ÎÒÃÇ·¢Ã÷£¬£¬£¬£¬£¬£¬£¬£¬¿ªÔ´Ä£×ӳ䵱Óû§Ê±¾­³£ÎÞ·¨ÎȹÌ×ñÕÕÖ¸Á£¬£¬£¬£¬£¬£¬£¬ÉõÖÁ»áËæÒâŲÓù¤¾ß£¬£¬£¬£¬£¬£¬£¬£¬µ¼Ö rollout Ìáǰʧ°Ü¡£¡£¡£¡£ ¡£¡£¡£¡£

ÔÚ¶àÂÖTool-Using AgentµÄѵÁ·ÖУ¬£¬£¬£¬£¬£¬£¬£¬reward²»ÔÙֻȡ¾öÓÚijһ´Î¹¤¾ßŲÓÃÊÇ·ñÀֳɣ¬£¬£¬£¬£¬£¬£¬£¬¶øÊÇÓÉÕû¶Î½»»¥trajectoryµÄ×îÖÕ״̬ͳһ¾öÒé¡£¡£¡£¡£ ¡£¡£¡£¡£ÕâÒâζ×Å£¬£¬£¬£¬£¬£¬£¬£¬Ö»ÒªÁ´Â·ÖÐÈκÎÒ»¸ö»·½Ú·ºÆðÎó²î£ºÒ»´ÎÓû§ÐÐΪÒì³£¡¢Ò»´Î¹¤¾ßÎóŲÓá¢Ò»´Î״̬ÌáǰÖÕÖ¹£¬£¬£¬£¬£¬£¬£¬£¬Õû¶ÎrolloutµÄreward¶¼¿ÉÄܱ»Ö±½Ó¹éÁã¡£¡£¡£¡£ ¡£¡£¡£¡£

´ÓЧ¹ûÉÏ¿´£¬£¬£¬£¬£¬£¬£¬£¬Agent¡¸Ê§°Ü¡¹ÁË£»£» £»£»£»£»µ«´ÓϵͳÄÚ²¿¿´£¬£¬£¬£¬£¬£¬£¬£¬Ê§°Ü²¢·×Æç¶¨À´×Ôagent policy×Ô¼º£¬£¬£¬£¬£¬£¬£¬£¬Ò²¿ÉÄÜÀ´×ÔÓÚÓû§Ä£×Ó×Ô¼ºµÄ²»ÎȹÌÐÔ¡£¡£¡£¡£ ¡£¡£¡£¡£

ÔÚÕæÊµÑµÁ·Àú³ÌÖУ¬£¬£¬£¬£¬£¬£¬£¬user modelÍùÍù²¢²»¿ÉʼÖÕÎȹ̵Ø×ñÕÕʹÃüÉ趨¡£¡£¡£¡£ ¡£¡£¡£¡£Ëü¿ÉÄÜÆ«ÀëÖ¸Áî¡¢ÎóŲÓù¤¾ß£¬£¬£¬£¬£¬£¬£¬£¬ÉõÖÁÔÚÒªº¦°ì·¨Ìáǰ¿¢Ê¶Ի°¡£¡£¡£¡£ ¡£¡£¡£¡£

ÕâЩÐÐΪ×Ô¼º²¢·Çagent¾öÒéµÄЧ¹û£¬£¬£¬£¬£¬£¬£¬£¬È´»áÖ±½Ó¾öÒé×îÖÕreward¡£¡£¡£¡£ ¡£¡£¡£¡£

ÓÚÊÇ£¬£¬£¬£¬£¬£¬£¬£¬ÇéÐξÍÄð³ÉAgentÔÚ¾Ö²¿¾öÒéÉÏÊÇ׼ȷµÄ£¬£¬£¬£¬£¬£¬£¬£¬µ«ÓÉÓÚÓû§ÐÐÎªÆ«ÒÆ£¬£¬£¬£¬£¬£¬£¬£¬×îÖÕÇéÐÎ״̬ʧ°Ü£¬£¬£¬£¬£¬£¬£¬£¬reward±»Í³Ò»ÅÐΪ0

´ÓÇ¿»¯Ñ§Ï°µÄÊӽǿ´£¬£¬£¬£¬£¬£¬£¬£¬Õâ×é³ÉÁËÑÏÖØµÄcredit assignment failure¡£¡£¡£¡£ ¡£¡£¡£¡£rewardÎÞ·¨Çø·Öʧ°ÜÊÂʵԴÓÚ agent policy£¬£¬£¬£¬£¬£¬£¬£¬ÕÕ¾ÉÀ´×Ôuser policyµÄÒì³£ÐÐΪ¡£¡£¡£¡£ ¡£¡£¡£¡£ÔÚÕâÖÖÌõ¼þÏ£¬£¬£¬£¬£¬£¬£¬£¬Ç¿»¯Ñ§Ï°²¢²»»á¡¸ÐÞÕý¡¹ÎÊÌ⣬£¬£¬£¬£¬£¬£¬£¬¶øÊÇ»áÒ»Ö±½«ÔëÉù·´ÏòÈö²¥µ½agentÉÏ£¬£¬£¬£¬£¬£¬£¬£¬×îÖÕÍÆ¶¯Õ½ÂÔ³¯×ŹýʧƫÏòÊÕÁ²¡£¡£¡£¡£ ¡£¡£¡£¡£

´ÓÕâ¸ö½Ç¶È¿´£¬£¬£¬£¬£¬£¬£¬£¬¶àÂÖAgentµÄѵÁ·Æ¿¾±£¬£¬£¬£¬£¬£¬£¬£¬²¢²»ÍêÈ«ÊÇËã·¨ÎÊÌ⣬£¬£¬£¬£¬£¬£¬£¬¶øÊÇÒ»¸öϵͳ½á¹¹ÎÊÌâ¡£¡£¡£¡£ ¡£¡£¡£¡£

»ùÓÚÕâÒ»Åжϣ¬£¬£¬£¬£¬£¬£¬£¬ÂÛÎIJ¢Ã»ÓмÌÐøÔÚÇ¿»¯Ñ§Ï°Ëã·¨²ãÃæµþ¼ÓÖØ´óÐÔ£¬£¬£¬£¬£¬£¬£¬£¬¶øÊÇÑ¡Ôñ´Ó¸üµ×²ãµÄѵÁ·Á÷³ÌÈëÊÖ£¬£¬£¬£¬£¬£¬£¬£¬ÖØÐ²ð½âagentÓëuserµÄ½ÇÉ«·Ö¹¤¡£¡£¡£¡£ ¡£¡£¡£¡£

EigenData²»¡¸ÌìÉú¸ü´ó¶¼¾Ý¡¹

ÈÃÊý¾Ý×Ô¼º½ø»¯

ÔÚ¶àÂÖTool-Using AgentµÄѵÁ·ÖУ¬£¬£¬£¬£¬£¬£¬£¬Êý¾ÝÎÊÌâÍùÍù±»¼ò»¯ÎªÒ»¸öÊýÄ¿ÎÊÌ⣺Êý¾Ý¹»²»·ó¶à¡¢ÁýÕÖ¹»²»·ó¹ã¡£¡£¡£¡£ ¡£¡£¡£¡£

µ«ÔÚÕæÊµlong-horizon½»»¥³¡¾°Ï£¬£¬£¬£¬£¬£¬£¬£¬Õâ¸ö¼ÙÉè²¢²»½¨Éè¡£¡£¡£¡£ ¡£¡£¡£¡£

´ó×Ú synthetic data ÔÚÎı¾²ãÃæ¿´ÆðÀ´ºÏÀí£¬£¬£¬£¬£¬£¬£¬£¬Âß¼­×ÔÇ¢¡¢¶Ô»°ÍêÕû£¬£¬£¬£¬£¬£¬£¬£¬µ«Ò»µ©ÕæÕýÖ´Ðй¤¾ßŲÓ㬣¬£¬£¬£¬£¬£¬£¬¾Í»á̻¶³ö¸ùÌìÐÔÎÊÌ⣺¹¤¾ß²ÎÊý²»Õýµ±¡¢×´Ì¬ÎÞ·¨µÖ´ï¡¢Ê¹ÃüÔÚÖÐ;½øÈë²»¿ÉÍê³ÉÇøÓò¡£¡£¡£¡£ ¡£¡£¡£¡£

ÕâÒâζ×Å£¬£¬£¬£¬£¬£¬£¬£¬Ä£×Ó²¢²»ÊÇÔÚ¡¸Ê§°ÜÖÐѧϰ¡¹£¬£¬£¬£¬£¬£¬£¬£¬¶øÊÇÔÚÓò»¿ÉÖ´ÐеĹ켣ѵÁ·×Ô¼º¡£¡£¡£¡£ ¡£¡£¡£¡£Òò´ËÔ­ÎÄÖÐEigenDataµÄÉè¼ÆÖØµã¹Ø×¢ÁËÔõÑù¹¹½¨Ò»¸ö¿É±Õ»·ÑÝ»¯µÄÊý¾ÝÌìÉúÀú³Ì£¬£¬£¬£¬£¬£¬£¬£¬¼´£º

ÌìÉúÊý¾Ý ¡ú ·¢Ã÷ʧ°Ü ¡ú ×Ô¶¯ÐÞÕýpromptÓëworkflow ¡ú ÔÙÌìÉú

EigenData²¢²»ÊǹŰåÒâÒåÉϵÄsynthetic data pipeline£¬£¬£¬£¬£¬£¬£¬£¬¶øÊÇÒ»¸öÄܹ»Æ¾Ö¤Ê§°Ü·´ÏìÒ»Á¬µü´úµÄ¶àÖÇÄÜϵһÇУ¬£¬£¬£¬£¬£¬£¬£¬ÍŽá×Ô¼ìÓë×ÔÐÞ¸´»úÖÆ£¬£¬£¬£¬£¬£¬£¬£¬Öð²½¹¹½¨³ö¸ßÖÊÁ¿µÄÊý¾ÝÜöÝÍ¡£¡£¡£¡£ ¡£¡£¡£¡£

ÔÚEigenDataµÄÊÂÇéÁ÷³ÌÖУ¬£¬£¬£¬£¬£¬£¬£¬Ã¿ÌõѵÁ·Ñù±¾¶¼±»ÒªÇó±ØÐèÖª×ãÒ»¸öÓ²ÐÔÌõ¼þ£ºÆä¶ÔÓ¦µÄ¹¤¾ßŲÓù켣¿£¿£¿ £¿£¿£¿£¿ÉÒÔ±»ÍêÕûÖ´ÐУ¬£¬£¬£¬£¬£¬£¬£¬²¢ÓÉverifierÔÚ´úÂë²ãÃæÑéÖ¤×îÖÕÇéÐÎ״̬¡£¡£¡£¡£ ¡£¡£¡£¡£

ÈôÊÇÖ´ÐÐʧ°Ü£¬£¬£¬£¬£¬£¬£¬£¬Ê§°ÜÐÅÏ¢»á±»»ØÁ÷£¬£¬£¬£¬£¬£¬£¬£¬ÓÃÓÚ×Ô¶¯ÐÞÕý prompt¡¢workflow ÒÔ¼°ÌìÉúÕ½ÂÔ×Ô¼º¡£¡£¡£¡£ ¡£¡£¡£¡£

ÕâʹµÃÊý¾ÝÂþÑܲ¢²»ÊÇÒ»´ÎÐÔÌìÉúµÄЧ¹û£¬£¬£¬£¬£¬£¬£¬£¬¶øÊÇ»áËæ×Åʧ°Ü·´ÏìÒ»Á¬Ïò¡¸¿ÉÖ´ÐÐÇøÓò¡¹ÊÕÁ²¡£¡£¡£¡£ ¡£¡£¡£¡£Í¨¹ý×Ô¶¯ÌìÉú¶àÂÖ¶Ô»°²¢Ö´ÐÐÕæÊµ¹¤¾ßŲÓ㬣¬£¬£¬£¬£¬£¬£¬Ã¿Ò»ÌõÊý¾ÝʵÀý¶¼»áÅäÌ×Ò»¸ö¡¸¿ÉÖ´ÐÐÑéÖ¤Æ÷¡¹£¬£¬£¬£¬£¬£¬£¬£¬Ê¹µÃ Agent ÐÐΪÊÇ·ñÀֳɿÉÒÔͨ¹ý´úÂëÖ±½ÓÅжϣ¬£¬£¬£¬£¬£¬£¬£¬Òò´ËÄܹ»°ü¹ÜÊý¾ÝÖÊÁ¿¡¸Ô½ÅÜÔ½ºÃ¡¹¡£¡£¡£¡£ ¡£¡£¡£¡£

´Óϵͳ½Ç¶È¿´£¬£¬£¬£¬£¬£¬£¬£¬Í¨¹ýÕâÒ»Ðж¯£¬£¬£¬£¬£¬£¬£¬£¬EigenDataÒ»Ö±ËõСÁËÄ£×Ó¿ÉÒÔѧϰµ½µÄÐÐΪ¿Õ¼ä£¬£¬£¬£¬£¬£¬£¬£¬Ê¹Æä¶ÔÆëÕæÊµÏµÍ³µÄ¿ÉÐн⼯¡£¡£¡£¡£ ¡£¡£¡£¡£ÕâÒ»²½°ü¹ÜÁËÄ£×ÓÔÚRL½éÈë֮ǰ£¬£¬£¬£¬£¬£¬£¬£¬Ã¿¸öreward¶¼¿ÉÒÔÕæÕý¶ÔÓ¦µ½Ò»¸öÒѾ­±»ÏµÍ³ÑéÖ¤ºóµÄЧ¹û£¬£¬£¬£¬£¬£¬£¬£¬Ê¹ÑµÁ·ÐźÅ×Ô¼ºÊÇ¿ÉÖ´ÐС¢¿ÉÑéÖ¤¡¢¿É¸´Ïֵġ£¡£¡£¡£ ¡£¡£¡£¡£

ÏÈѵÓû§Ä£×Ó£¬£¬£¬£¬£¬£¬£¬£¬ÔÙѵAgent

¼´±ãѵÁ·Êý¾Ý×Ô¼ºÊÇ¿ÉÖ´Ðе쬣¬£¬£¬£¬£¬£¬£¬¶àÂÖ Agent µÄѵÁ·ÈÔÈ»¿ÉÄÜʧ°Ü¡£¡£¡£¡£ ¡£¡£¡£¡£

Ôµ¹ÊÔ­ÓÉÔÚÓÚ£¬£¬£¬£¬£¬£¬£¬£¬ÔÚinteractive agent³¡¾°ÖУ¬£¬£¬£¬£¬£¬£¬£¬Óû§Ä£×Ó×Ô¼º¾ÍÊÇϵͳµÄÒ»²¿·Ö¡£¡£¡£¡£ ¡£¡£¡£¡£

ÈôÊÇuser policy±£´æÆ¯ÒÆ»ò²»ÎȹÌÐÔ£¬£¬£¬£¬£¬£¬£¬£¬¼´±ã agent µÄ¾Ö²¿¾öÒéÊÇ׼ȷµÄ£¬£¬£¬£¬£¬£¬£¬£¬Õû¶Î trajectory ÈÔ¿ÉÄÜÓÉÓÚÓû§ÐÐΪÒì³£¶øÊ§°Ü£¬£¬£¬£¬£¬£¬£¬£¬×îÖÕ reward ±»Í³Ò»¹éÁã¡£¡£¡£¡£ ¡£¡£¡£¡£

»ùÓÚÕâÒ»ÊìϤ£¬£¬£¬£¬£¬£¬£¬£¬Ñо¿ÕßÃǽ«ÑµÁ·Á÷³Ì²ð·ÖΪÁ½²½£º

Ê×ÏÈ£¬£¬£¬£¬£¬£¬£¬£¬Ê¹ÓÃEigenDataÌìÉúµÄ¿ÉÖ´ÐжԻ°Êý¾Ý£¬£¬£¬£¬£¬£¬£¬£¬¶Ôuser model¾ÙÐÐSFT΢µ÷£¬£¬£¬£¬£¬£¬£¬£¬Ê¹ÆäÐÐΪÎȹ̡¢¿É¿Ø£¬£¬£¬£¬£¬£¬£¬£¬²¢ÓëʹÃüÉ趨¶ÔÆë£»£» £»£»£»£»

ÔÚÓû§²à²»ÔÙ³ÉΪÖ÷ÒªÔëÉùÔ´Ö®ºó£¬£¬£¬£¬£¬£¬£¬£¬²ÅÒýÈëÇ¿»¯Ñ§Ï°ÓÅ»¯agent policy¡£¡£¡£¡£ ¡£¡£¡£¡£

ÕâÒ»²ð·Ö²¢²»ÊÇÌØÁíÍ⹤³ÌÖØÆ¯ºó£¬£¬£¬£¬£¬£¬£¬£¬¶øÊÇÒ»¸öϵͳ¼¶Ç°ÖÃÌõ¼þ¡£¡£¡£¡£ ¡£¡£¡£¡£Ëü´Ó»ù´¡ÉÏïÔÌ­ÁË reward µÄ»ìÔÓȪԴ£¬£¬£¬£¬£¬£¬£¬£¬Ê¹Ç¿»¯Ñ§Ï°²»ÔÙÆµÈÔ´¦·Ö¡¸×¼È·µ«±»Óû§ÐÐÎªÆÆËðµÄ¾öÒ项£¬£¬£¬£¬£¬£¬£¬£¬ÑµÁ·ÇúÏßÒ²Òò´Ë±äµÃÎȹ̡¢¿ÉÕ¹Íû¡£¡£¡£¡£ ¡£¡£¡£¡£

ÓḿÉÖ´ÐÐЧ¹û¡¹Ìæ»»Ö÷¹Û½±Àø

ÔÚÇ¿»¯Ñ§Ï°½×¶Î£¬£¬£¬£¬£¬£¬£¬£¬¸ÃÒªÁì²»ÔÙÒÀÀµÄ£ºýµÄreward model£¬£¬£¬£¬£¬£¬£¬£¬¶øÊÇÓÃʹÃü×Ô´øµÄÑéÖ¤º¯Êý£¨verifier£©Ö±½Ó¼ì²é×îÖÕÇéÐÎ״̬£¬£¬£¬£¬£¬£¬£¬£¬ÊµÏÖ¡¸¶Ô / ´í¡¹µÄ¿ÉÖ´ÐС¢¿ÉÉ󼯽±ÀøÐźš£¡£¡£¡£ ¡£¡£¡£¡£

ÔÚ´Ë»ù´¡ÉÏ£¬£¬£¬£¬£¬£¬£¬£¬ÒýÈëGRPOµÄgroup-relative advantage£ºÕë¶ÔͳһʹÃü²ÉÑù¶àÌõtrajectory£¬£¬£¬£¬£¬£¬£¬£¬¾ÙÐÐ×éÄÚÏà¶ÔÓÅÊÆÑ§Ï°£¬£¬£¬£¬£¬£¬£¬£¬ÒÔ½µµÍlong-horizon½»»¥µ¼Öµĸ߷½²îÓë²»ÎȹÌÐÔ¡£¡£¡£¡£ ¡£¡£¡£¡£

ͬʱʹÓÃdynamic filteringÌÞ³ý¡¸È«¶Ô/È«´í¡¹µÄµÍÐÅÏ¢Ñù±¾£¬£¬£¬£¬£¬£¬£¬£¬½«ÑµÁ·Ô¤Ë㼯ÖÐÓÚ¾ßÓÐÇø·Ö¶ÈµÄʹÃü×Ó¼¯¡£¡£¡£¡£ ¡£¡£¡£¡£

ÔÚÕâЩÉè¼ÆµÄÅäÏàÖúÓÃÏ£¬£¬£¬£¬£¬£¬£¬£¬RLÐźŸüÇå½à¡¢¸üÎȹ̣¬£¬£¬£¬£¬£¬£¬£¬ÑµÁ·Àú³ÌÒ²¸ü²»Ò×·ºÆðÕ½ÂÔÆ¯ÒÆ¡£¡£¡£¡£ ¡£¡£¡£¡£

ʵÑéЧ¹û

¿ªÔ´Ä£×ÓѵÁ·ÖÁ¿¿½ü¹Ø±ÕÄ£×ÓË®×¼

ΪÁËÑéÖ¤ÕâÒ»Ì×ϵͳ¼¶ÑµÁ··¶Ê½ÔÚÕæÊµ½»»¥³¡¾°ÖеÄÓÐÓÃÐÔ£¬£¬£¬£¬£¬£¬£¬£¬Ñо¿ÕßÔÚ¦Ó?-benchµÄÈý¸öÕæÊµ¹¤¾ßʹÃü£¨Airline / Retail / Telecom£©ÉϾÙÐÐÁËϵͳÆÀ¹À¡£¡£¡£¡£ ¡£¡£¡£¡£ÆÀ¹À½ÓÄÉpass?Ö¸±ê£¬£¬£¬£¬£¬£¬£¬£¬¼´ÒªÇóAgentÔÚÒ»´ÎÍêÕû¶àÂÖ½»»¥ÖÐÀÖ³ÉÍê³ÉʹÃü£¬£¬£¬£¬£¬£¬£¬£¬ÕâÒ»Ö¸±êÄܹ»¸üÖ±½Ó·´Ó¦ Agent ÔÚ long-horizon ³¡¾°ÏµÄÎȹÌÐÔÓë¿É¿¿ÐÔ¡£¡£¡£¡£ ¡£¡£¡£¡£

Ч¹ûÏÔʾ£¬£¬£¬£¬£¬£¬£¬£¬ÐÔÄÜÌáÉý²¢·ÇÎÞÒ⣬£¬£¬£¬£¬£¬£¬£¬¶øÊÇÔÚ¶à¸ö³¡¾°ÖÐÎȹ̷ºÆð¡£¡£¡£¡£ ¡£¡£¡£¡£

ÔÚ¹æÔò×îÖØ´óµÄTelecom³¡¾°ÖУ¬£¬£¬£¬£¬£¬£¬£¬Qwen3-235B-A22B-2507¾­SFT + RLѵÁ·ºó£¬£¬£¬£¬£¬£¬£¬£¬pass?ÌáÉýÖÁ98.3%£¬£¬£¬£¬£¬£¬£¬£¬½øÈëÄ¿½ñ¹ûÕæÐ§¹ûµÄ×îÇ¿Ìݶӣ»£» £»£»£»£»

ÔÚAirline³¡¾°ÖУ¬£¬£¬£¬£¬£¬£¬£¬Í³Ò»Ä£×ÓµÖ´ï73.0% pass?£¬£¬£¬£¬£¬£¬£¬£¬ÕûÌåÌåÏÖÒÑÓëÖ÷Á÷±ÕԴϵͳ¶ÔÆë¡£¡£¡£¡£ ¡£¡£¡£¡£

¸üÒªº¦µÄÊÇ£¬£¬£¬£¬£¬£¬£¬£¬ÔÚÈýÓò»ìÏýѵÁ·ÉèÖÃÏ£¬£¬£¬£¬£¬£¬£¬£¬Ò»¸öÄ£×Óͬʱѧϰ¶à¸ö¹¤¾ßÇéÐΣ¬£¬£¬£¬£¬£¬£¬£¬×îÖÕÈÔÄܼá³Ö81.3% µÄƽ¾ù pass?£¬£¬£¬£¬£¬£¬£¬£¬Åú×¢¸ÃÒªÁìѧµ½µÄ²¢·Ç¼òµ¥³¡¾°Ïµġ¸Í¶ÆõÕ½ÂÔ¡¹£¬£¬£¬£¬£¬£¬£¬£¬¶øÊǸü¾ßͨÓÃÐ﵀ tool-using ÄÜÁ¦¡£¡£¡£¡£ ¡£¡£¡£¡£

½øÒ»²½µÄÏûÈÚʵÑéÕ¹ÏÖÁËÕâЩÌáÉýµÄȪԴ¡£¡£¡£¡£ ¡£¡£¡£¡£

Ò»µ©ÒƳývalidation / verifier»òÊý¾Ý×Ô½ø»¯»úÖÆ£¬£¬£¬£¬£¬£¬£¬£¬SFT ½×¶ÎµÄÐÔÄܱ㷺ÆðÏÔ×ÅϽµ£¬£¬£¬£¬£¬£¬£¬£¬ËµÃ÷Êý¾ÝµÄ¿ÉÖ´ÐÐÐÔÓë¶àÑùÐÔÊÇÄÜÁ¦ÐγɵĻù´¡£¡£¡£¡£ ¡£¡£¡£¡£»£» £»£»£»£»¶øÈôÊÇÔÚδ¶ÔÓû§Ä£×Ó¾ÙÐÐÎȹ̻¯Ô¤ÑµÁ·µÄÇéÐÎÏÂÖ±½ÓÒýÈëÇ¿»¯Ñ§Ï°£¬£¬£¬£¬£¬£¬£¬£¬ÕûÌåÐÔÄÜ·´¶ø»áÍË»¯¡£¡£¡£¡£ ¡£¡£¡£¡£ÕâһЧ¹ûÅú×¢£¬£¬£¬£¬£¬£¬£¬£¬Ö»ÓÐÔÚÓû§ÐÐΪ±»ÓÐÓÿØÖƵÄÌõ¼þÏ£¬£¬£¬£¬£¬£¬£¬£¬Ç¿»¯Ñ§Ï°²Å»ªÒ»Á¬´øÀ´ÕýÏòÔöÒæ¡£¡£¡£¡£ ¡£¡£¡£¡£

¿ÉÖ´ÐÐѵÁ·ÐźŲ¢²»ÊÇÒ»¸ö¡¸½õÉÏÌí»¨¡¹µÄ¼¼ÇÉ£¬£¬£¬£¬£¬£¬£¬£¬¶øÊÇÒ»ÌõÃ÷È·µÄϵͳ·Ö½çÏß¡£¡£¡£¡£ ¡£¡£¡£¡£

µ± Tool-Using Agent ½øÈëÕæÊµ¶àÂÖ½»»¥£¬£¬£¬£¬£¬£¬£¬£¬ÎÊÌâ²»ÔÙÖ»ÊÇ¡¸Ç¿»¯Ñ§Ï°»¹Äܲ»¿ÉÊÕÁ²¡¹£¬£¬£¬£¬£¬£¬£¬£¬¶øÊÇѵÁ·ÐźÅ×Ô¼ºÊÇ·ñ¾ß±¸¹¤³ÌÒâÒ壺ËüÊÇ·ñ¿ÉÖ´ÐС¢¿É¹éÒò¡¢¿ÉÑéÖ¤£¬£¬£¬£¬£¬£¬£¬£¬ÊÇ·ñÕæÕý¶ÔÓ¦µ½Ò»¸ö¿É¸´ÏÖµÄϵͳЧ¹û¡£¡£¡£¡£ ¡£¡£¡£¡£ÕâÕýÊÇEigenData½éÈëµÄλÖᣡ£¡£¡£ ¡£¡£¡£¡£

ͨ¹ý½«Êý¾ÝÌìÉú¡¢¹¤¾ßÖ´ÐÐÓëverifierУÑéͳһ½øÒ»¸ö±Õ»·ÏµÍ³£¬£¬£¬£¬£¬£¬£¬£¬EigenData²»µ«ÊÇΪRLÌṩÁË¡¸¸üÇå½àµÄreward¡¹£¬£¬£¬£¬£¬£¬£¬£¬¶øÊÇÖØÐ½ç˵ÁËʲôÑùµÄѵÁ·ÐźŲÅÖµµÃ±»Ç¿»¯Ñ§Ï°·Å´ó¡£¡£¡£¡£ ¡£¡£¡£¡£ÔÚÕâÒ»Ìõ¼þÏ£¬£¬£¬£¬£¬£¬£¬£¬GRPO¡¢dynamic filteringµÈÓÅ»¯Õ½ÂԲŵÚÒ»´ÎÓµÓÐÇåÎú¡¢Îȹ̵Ä×÷Óù¤¾ß¡£¡£¡£¡£ ¡£¡£¡£¡£

ÂÛÎĸø³öµÄÅжϱê×¼×ÅʵºÜÊÇÖ±½Ó£ºÈôÊÇÒ»¸ö¶àÂÖAgentµÄѵÁ·Á÷³ÌÎÞ·¨Ã÷È·»Ø¸²¡¸reward ¾¿¾¹ÔÚ½±ÀøË­¡¢Ê§°ÜÊÂʵÓÉË­µ¼Ö¡¢Í³Ò»Ê¹ÃüÏÂÄÄÌõ¹ì¼£¸üºÃ¡¹£¬£¬£¬£¬£¬£¬£¬£¬ÄÇËüÔÚ¹¤³ÌÉÏÈÔÍ£ÁôÔÚ¡¸¿´ÆðÀ´ÄÜÅÜ¡¹µÄ workflow£¬£¬£¬£¬£¬£¬£¬£¬¶ø²»ÊÇ¡¸¿ÉÒÔÒ»Á¬ÓÅ»¯¡¹µÄsystem¡£¡£¡£¡£ ¡£¡£¡£¡£

´ÓÕâ¸ö½Ç¶È¿´£¬£¬£¬£¬£¬£¬£¬£¬ÑµÁ·ÖзºÆðµÄperformance oscillation¡¢reward ±»Òì³£Óû§ÐÐΪÇåÁã¡¢RL ·´¶ø´øÀ´ÍË»¯£¬£¬£¬£¬£¬£¬£¬£¬²¢²»ÊÇʵÏÖϸ½ÚÉϵÄ覴㬣¬£¬£¬£¬£¬£¬£¬¶øÊÇѵÁ·ÐźÅÉÐδ±»ÏµÍ³ÐԽṹµÄÒ»¶¨Ð§¹û¡£¡£¡£¡£ ¡£¡£¡£¡£

ÕâÏîÊÂÇéµÄ½¹µãТ˳£¬£¬£¬£¬£¬£¬£¬£¬²¢²»ÔÚÓÚÌá³öÒ»ÖÖеÄRL¼¼ÇÉ£¬£¬£¬£¬£¬£¬£¬£¬¶øÔÚÓÚͨ¹ýEigenData½«¶àÂÖAgentµÄpost-trainingÍÆÏòÒ»¸öÐµĹ¤³Ì·¶Ê½£º

µ±ÑµÁ·ÐźÅÏȱ»½á¹¹³É¿ÉÖ´ÐС¢¿É¹éÒò¡¢¿ÉÑéÖ¤µÄϵͳ¹¤¾ßʱ£¬£¬£¬£¬£¬£¬£¬£¬Ç¿»¯Ñ§Ï°²ÅÕæÕý³ÉΪһÖֿɿصÄϵͳÓÅ»¯£»£» £»£»£»£»ÔÚ´Ë֮ǰ£¬£¬£¬£¬£¬£¬£¬£¬ÔÙ¶àµÄ rollout ºÍ¸ü´óµÄÄ£×Ó£¬£¬£¬£¬£¬£¬£¬£¬Ò²Ö´ÙÇÔÚÔëÉùÖ®Éϵþ¼ÓÅÌËã¡£¡£¡£¡£ ¡£¡£¡£¡£

²Î¿¼×ÊÁÏ£º

https://arxiv.°²»Õ½­»´Æû³µ²¿¼þÓÐÏÞ¹«Ë¾org/abs/2601.22607