¶¥¼âÄ£×ÓÀë¡°¿ÆÑ§¼Ò¡±»¹²îµÃÔ¶£¿£¿£¿£¿£¿£¿£¿AI4Sؽ´ýÂõÏò2.0ʱ´ú
2026-02-26 20:08:01

»úе֮ÐÄÐû²¼

µ±ÔÆÄÏµá³ØÂÌɫʳÎïÓÐÏÞ¹«Ë¾Ç°£¬£¬£¬£¬£¬¿ÆÑ§ÖÇÄÜ£¨AI for Science£©±»³ÆÖ®ÎªÈ˹¤ÖÇÄÜµÄ ¡°»Ê¹Ú¡±£¬£¬£¬£¬£¬ÒÔ AlphaFold Ϊ´ú±íµÄ AI for Science£¨AI4S£©ÊÖÒÕÔÚÂѰ×ÖÊÕÛµþ¡¢ÆøÏóÕ¹ÍûµÈÌØ¶¨ÁìÓòÈ¡µÃÁËÀï³Ì±®Ê½³É¼¨£¬£¬£¬£¬£¬µ«½üÆÚ¡¶Nature¡·½ÒÏþµÄÑо¿Ö¸³ö£¬£¬£¬£¬£¬Ì«¹ýÒÀÀµÏÖÓÐÉî¶Èѧϰģ×Ó¿ÉÄܾÖÏÞÐÂ֪ʶµÄ̽Ë÷½çÏߣ¬£¬£¬£¬£¬ÉõÖÁÔÚijÖÖˮƽÉÏ×è°­Á¢Òì¡£¡£¡£¡£¡£¡£¡£

Ò»ÏîÀ´×ÔÉϺ£È˹¤ÖÇÄÜʵÑéÊÒ£¨ÉϺ£ AI Lab£©µÄϵͳÐÔÆÀ¹À¢Ù½øÒ»²½Õ¹ÏÖÁËÄ¿½ñÇ°ÑØÄ£×ӵĶ̰塣¡£¡£¡£¡£¡£¡£À´×Ô 10 ¸ö²î±ð¿ÆÑ§ÁìÓòµÄ 100 λ¿ÆÑ§¼ÒΪģ×Ó¹¹½¨ÁËÆÀ²âÎÊÌ⣬£¬£¬£¬£¬Ð§¹ûÏÔʾ£ºÇ°ÑØÄ£×ÓÔÚͨÓÿÆÑ§ÍÆÀíʹÃüÖе÷ֿɴï 50 ·Ö£¨Âú·Ö 100£©£¬£¬£¬£¬£¬µ«ÔÚÖÖÖÖ×¨ÒµÍÆÀíʹÃü£¨ÈçרÏîÎÄÏ×¼ìË÷¡¢ÏêϸʵÑ鼯»®Éè¼Æ£©ÖУ¬£¬£¬£¬£¬µÃ·ÖÖè½µÖÁ 15-30 ·Ö¡£¡£¡£¡£¡£¡£¡£

¡°ÎÒÃÇÒÑÉí´¦ ¡°Í¨ÓÃÈ˹¤ÖÇÄÜ¡±£¨AGI£©Ç°Ï¦£¬£¬£¬£¬£¬µ«ÈÔÃæÁÙÖ÷Òª»·½ÚµÄȱʧ ¡ª¡ª ͨרÈںϵÄÖÇÄÜ¡£¡£¡£¡£¡£¡£¡£ÎÒÃÇØ½ÐèÍÆ¶¯¿ÆÑ§ÖÇÄÜ´Ó 1.0 Ïò 2.0 µü´ú£¬£¬£¬£¬£¬¼´´Ó AI4S ÂõÏò AGI4S¡£¡£¡£¡£¡£¡£¡£¡± ÈÕǰ£¬£¬£¬£¬£¬ÉϺ£È˹¤ÖÇÄÜʵÑéÊÒÖ÷ÈΡ¢Ê×ϯ¿ÆÑ§¼ÒÖܲ®ÎÄÔÚµÚËÄÊ®½ìÈ˹¤ÖÇÄÜЭ»áÄê»á£¨AAAI 2026£©½ÒÏþÌØÑû±¨¸æÊ±Ìá³ö£¬£¬£¬£¬£¬¿ÆÑ§·¢Ã÷ÊÇ AI µÄÏÂÒ»¸öÇ°ÑØÕóµØ ¡ª¡ª Ëü¼ÈÊÇÍÆÀíÖÇÄܵÄ×îÖÕÊÔÁ¶³¡£¬£¬£¬£¬£¬Ò²ÊÇ ¡°Í¨×¨ÈÚºÏ AGI¡± µÄÑéÖ¤Îę̀¡£¡£¡£¡£¡£¡£¡£Èô AGI = ͨרÈںϣ¨Specialized Generalist£©£¬£¬£¬£¬£¬Ôò¿ÉÉî¶Èרҵ»¯Í¨ÓÃÄ£×Ó£¨Specializable Generalist£©ÊÇʵÏÖ AGI µÄ¿ÉÐз¾¶¡£¡£¡£¡£¡£¡£¡£

³ýÁË·ÖÏíÇ°ÑØ¿´·¨£¬£¬£¬£¬£¬Öܲ®ÎÄ»¹ÏêϸÏÈÈÝÁËÉϺ£ AI ʵÑéÊÒ½üÄêÀ´¿ªÕ¹µÄÇ°ÑØÌ½Ë÷Óëʵ¼ù£¬£¬£¬£¬£¬°üÀ¨Çý¶¯ ¡°Í¨×¨Èںϡ± Éú³¤µÄÊÖÒռܹ¹ ¡ª¡ª¡°ÖÇÕß¡±SAGE£¨Synergistic Architecture for Generalizable Experts£©£¬£¬£¬£¬£¬Æä°üÀ¨»ù´¡¡¢ÈÚºÏÓë½ø»¯Èý¸öÌõÀí£¬£¬£¬£¬£¬²¢¿ÉË«ÏòÑ­»·ÊµÏÖȫջ½ø»¯ £» £»£»£»£»£»Ö§³Ö AGI4S ̽Ë÷µÄÁ½´ó»ù´¡ÉèÊ©¡°ÊéÉú¡±¿ÆÑ§¶àģ̬´óÄ£×Ó Intern-S1¡¢¡°ÊéÉú¡±¿ÆÑ§·¢Ã÷ƽ̨ Intern-Discovery ¼°Ò»ÏµÁÐÏà¹Ø½×¶ÎÐÔÏ£Íû¡£¡£¡£¡£¡£¡£¡£

Ñݽ²×îºó£¬£¬£¬£¬£¬Öܲ®ÎÄÏò»á³¡ÄÚÍâµÄ¹ÛÖÚ·¢³öÐж¯ÕÙ»½£º¼Ü¹¹ÒѾ­Í£µ±£¬£¬£¬£¬£¬µ«»­¾íÈÔ´æ´óƬÁô°×£¬£¬£¬£¬£¬ÆÚ´ýÓë¸ü¶àÙÉÐÐÕß¹²ÍØÀ¶Í¼£¡

ÒÔÏÂΪ±¨¸æÈ«ÎÄ£¬£¬£¬£¬£¬ÂÔÓÐÐÞ¶©¡£¡£¡£¡£¡£¡£¡£

ÑݽøÔ¤ÅУº´Ó ANI µ½ AGI µÄÀúÊ·¿çÔ½

È˹¤ÖÇÄܵÄÉú³¤Àú³Ì²¢·ÇÏßÐԶѵþ£¬£¬£¬£¬£¬¶øÊÇ·ºÆð³öÏÔ×ŵĽ׶ÎÐÔԾǨ¡£¡£¡£¡£¡£¡£¡ £» £»£»£»£»£»ØÊ× AI Éú³¤µÄÀúÊ·×ø±ê£¬£¬£¬£¬£¬ÓÐÖúÓÚÎÒÃÇÀåÇåÄ¿½ñËù´¦µÄλÖü°Î´À´µÄÆ«Ïò¡£¡£¡£¡£¡£¡£¡£

ÔçÔÚ 1996 ÄêÉæ×ã AI Ñо¿Ö®³õ£¬£¬£¬£¬£¬ÎÒ±ã×îÏÈ˼Ë÷ÖÇÄܵÄʵÖÊ¡£¡£¡£¡£¡£¡£¡£ÌØÊâÊÇÔÚµ£µ± IBM È˹¤ÖÇÄÜ»ù´¡Ñо¿ÔºÔººã¾Ã¼ä£¬£¬£¬£¬£¬Ê×´ÎÌá³öÁËͨÍùͨÓÃÈ˹¤ÖÇÄÜ£¨AGI£©µÄÕ½ÂÔõ辶ͼ£¬£¬£¬£¬£¬Ã÷È·½ç¶¨ÁË AI Éú³¤µÄÈý¸öÒªº¦½×¶Î£ºANI£¨ÏÁÒåÈ˹¤ÖÇÄÜ£©¡¢ABI£¨¹ãÒåÈ˹¤ÖÇÄÜ£©Óë AGI£¬£¬£¬£¬£¬²¢¸ø³öÁ˸÷×ÔÃ÷È·½ç˵¡£¡£¡£¡£¡£¡£¡£

ÎÒÆäʱµÄÅжÏÊÇ ANI ÔÚ 2016 ÄêÒÑÇ÷ÓÚ³ÉÊ죬£¬£¬£¬£¬¶øÍ¨Íù AGI µÄ±Ø¾­Ö®Â·²¢·ÇÖ±½ÓԾǨ£¬£¬£¬£¬£¬¶øÊDZØÐèÂÊÏÈʵÏ־߱¸¿çÁìÓò·º»¯ÄÜÁ¦µÄ ABI¡£¡£¡£¡£¡£¡£¡£ÎÒÃÇÒÔΪÕâÒ»¿çÔ½ÐèÒªÊÖÒÕ·¶Ê½µÄ¸ùÌìÐÔÀå¸ï£¬£¬£¬£¬£¬×îÉÙ°üÀ¨Èý¸ö·½Ã棺¼´´ÓÓмàÊÓѧϰתÏò×Ô¼àÊÓѧϰ£¬£¬£¬£¬£¬´ÓÈËÀàÖ§½âʹÃü¼¶ÁªÊ½ÏµÍ³×ªÏò¶Ëµ½¶Ë¼Ü¹¹£¬£¬£¬£¬£¬´ÓÅбðʽ¹¤¾ß½ø»¯ÎªÌìÉúʽÖúÊÖ¡£¡£¡£¡£¡£¡£¡£

ÁùÄê¶àºó ChatGPT µÄÎÊÊÀ£¬£¬£¬£¬£¬µÚÒ»´ÎÑéÖ¤ÁËÈ˹¤ÖÇÄÜϵͳÔÚÒÔÉÏÈý·½ÃæµÄͬʱ¸æ¿¢£¬£¬£¬£¬£¬ÊµÖÊÉÏÐû¸æÁË ABI ½×¶ÎµÄµ½À´¡£¡£¡£¡£¡£¡£¡£ÕâÒ»ÀúÊ·ÐÔÍ»ÆÆÑéÖ¤Á˹æÄ£¹æÔò£¨Scaling Law£©µÄÓÐÓÃÐÔ ¡ª¡ª ¼´Í¨¹ýÀ©´ó Transformer ¼Ü¹¹²¢½« ¡°ÏÂÒ»¸ö´ÊÕ¹Íû¡± ×÷ΪÓÅ»¯Ä¿µÄ£¬£¬£¬£¬£¬ÈËÀàÊ×´ÎʵÏÖÁ˶ÔÌìÏÂ֪ʶµÄѹËõ¡£¡£¡£¡£¡£¡£¡£ÖµµÃÒ»ÌáµÄÊÇ£¬£¬£¬£¬£¬ÎÒºÍÍŶÓÔçÔÚ 2016 ÄêÌá³öµÄ¹ØÓÚ ¡°¶àÍ·×Ô×¢ÖØÁ¦¡± »úÖÆµÄÑо¿£¬£¬£¬£¬£¬×÷Ϊ ¡°ÓëÏÂÓÎʹÃüÎ޹ء±£¨Ò²¾ÍÊÇ ¡°Ô¤ÑµÁ·¡±£©µÄ×ÔÈ»ÓïÑÔ³¤ÉÏÏÂÎÄѹËõ±íÕ÷µÄÊ×ÅúЧ¹ûÖ®Ò»£¬£¬£¬£¬£¬±»¿ª´´Ð﵀ Transformer ÂÛÎÄÒýÓÃÓëÈϿɢÚ£¬£¬£¬£¬£¬ÎªÕâһԤѵÁ·Ê±´úµÄѹËõÖÇÄܵÓÚ¨ÁËÖ÷ÒªµÄÀíÂÛ»ùʯ¡£¡£¡£¡£¡£¡£¡£

ÖØ·Ãõ辶ͼ£¨2016 Ä꣩£ºÍ¨Íù AGI ֮·

Õ½ÂÔ·¾¶£ºÍ¨×¨ÈÚºÏÓë¿ÆÑ§·¢Ã÷µÄ×îÖÕÊÔÁ¶

Ëæ×Å Scaling Law ¸¶ÓëÁË´óÓïÑÔÄ£×ÓÆÕ±éµÄ·º»¯ÄÜÁ¦£¨ABI£©£¬£¬£¬£¬£¬ÔÚ 2023 ÄêÍ·ÎÒÃÇÌá³öÁËÒ»¸öÒªº¦µÄÕ½ÂÔÉèÎÊ£ºÍ¨Íù AGI µÄÏÂÒ»²½£¬£¬£¬£¬£¬½ö½öÊÇÅÌËãÁ¿µÄ¶ÑµþÂ𣿣¿£¿£¿£¿£¿£¿¶ÔÕâЩÉèÎʵÄ˼Ë÷´ÙʹÎÒÔÚ 2023 ÄêÌá³öÁË¡°Í¨×¨Èںϡ± ·¾¶¡£¡£¡£¡£¡£¡£¡£½¹µãÍ·ÄÔÊÇÔõÑù¶¯Ì¬ÊµÑéÈÚºÏÈËÀàÈÏ֪ͷÄÔµÄϵͳ 1 ºÍϵͳ 2£¬£¬£¬£¬£¬ÒÔÓ¦¶ÔÖÖÖÖÏÖʵÌìϵÄʹÃü¡£¡£¡£¡£¡£¡£¡£

ÖØÐ½ç˵ AGI ֮·

ÒÑÍù 70 Äê AI µÄÉú³¤ºã¾ÃÔÚ ¡°×¨ÒµÐÔ¡± Óë ¡°Í¨ÓÃÐÔ¡± Á½¸öά¶ÈÉÏ»®·ÖÏ£Íû¡£¡£¡£¡£¡£¡£¡£ÒÔ AlphaFold Ϊ´ú±íµÄÔçÆÚϵͳÊǼ«Ö嵀 ¡°×¨¼Ò¡±£¬£¬£¬£¬£¬ÔÚÌØ¶¨ÁìÓòÓâÔ½ÈËÀàȴȱ·¦Ç¨áãÄÜÁ¦ £» £»£»£»£»£»¶øÄ¿½ñµÄ´óÓïÑÔÄ£×ÓÔòÊDz©ÎŹãʶµÄ ¡°Í¨²Å¡±£¬£¬£¬£¬£¬Ëä¾ß¹ã¶Èµ«ÔÚ´¦Öóͷ£ÖØ´óרҵʹÃüʱÍùÍùÄÑÒÔÆó¼°×¨¼ÒÉî¶ÈºÍȱʧҪº¦Ï¸½Ú¡£¡£¡£¡£¡£¡£¡£ÕæÕýµÄ AGI ±ØÐèÍ»ÆÆÕâÖÖ¶þÔª¶ÔÁ¢£¬£¬£¬£¬£¬¹¹½¨Ò»ÖÖÄܹ»¶¯Ì¬ÈÚºÏ ¡°ÏµÍ³ 1¡±£¨Ö±¾õʽ¿ì˼Ë÷£©Óë ¡°ÏµÍ³ 2¡±£¨Âß¼­Ê½Âý˼Ë÷£©µÄÖÇÄܼܹ¹ ¡ª¡ª ¼´ÔÚ¼á³ÖͨÓÃÈÏÖª»ù×ùµÄͬʱ£¬£¬£¬£¬£¬Äܹ»ÔÚí§ÒâÌØ¶¨Ê¹ÃüÉÏͨ¹ýÒ»Á¬Ñ§Ï°ÓëÉî¶ÈÍÆÀíʵÏÖר¼Ò¼¶µÄר¾«£¨ÐðÊöÕâһ˼Ð÷ϵͳµÄ̬¶ÈÂÛÎÄÒÑÓÚ 2024 ÄêÔÚ ArXiv ÉϽÒÏþ£©¢Û¡£¡£¡£¡£¡£¡£¡£

2024 Äêβ OpenAI o1 Óë 2025 ÄêÍ· DeepSeek-R1 µÄ·ºÆð£¬£¬£¬£¬£¬Í¨¹ýÔÚ´óÄ£×ÓÖ®ÉÏÓ¦ÓÃÇ¿»¯Ñ§Ï°ÏÔÖøÌáÉýÂß¼­ÍÆÀíÄÜÁ¦£¬£¬£¬£¬£¬ÓÐÁ¦µØÑéÖ¤Á˹ØÓÚ ¡°Í¨×¨Èںϡ± ·¾¶Ô¤ÅеÄ׼ȷÐÔ¡£¡£¡£¡£¡£¡£¡£2025 Äê 10 Ô£¬£¬£¬£¬£¬Ô¼ÊéÑÇ?±¾¼ª°Â½ÌÊÚµÈÈËÌá³öÁË AGI µÄ½ç˵£¬£¬£¬£¬£¬½«ÆäÆÊÎöΪʮÖÖ½¹µãͨÓÃÄÜÁ¦ÒÔ¼°ÖÚ¶àÏÁÒåµÄרҵÄÜÁ¦¡£¡£¡£¡£¡£¡£¡£ÈôÄÜÖÜÈ«¸æ¿¢ÕâЩÄÜÁ¦£¬£¬£¬£¬£¬¼´Òâζ×ÅʵÏÖÁË AGI¡£¡£¡£¡£¡£¡£¡£ÕâÒ»½ç˵ÓëÎÒÃÇ¡°Í¨×¨ÈÚºÏÊÇͨÍù AGI µÄÕ½ÂÔ·¾¶¡±µÄ¿´·¨¸ß¶ÈÎÇºÏ ¡ª¡ª ÕâÅú×¢¸Ã·¾¶ÕýÈÕÒæ³ÉΪÕû¸öѧÊõÉçÇøµÄÆÕ±é¹²Ê¶¡£¡£¡£¡£¡£¡£¡£

¿ÆÑ§·¢Ã÷£ºÍÆÀíÖÇÄܵÄ×îÖÕÇ°ÑØ

ÏÂÒ»¸öÇ°ÑØÁìÓòÊÇʲô£¿£¿£¿£¿£¿£¿£¿ÎÒÒÔΪÊÇ¿ÆÑ§·¢Ã÷£¨Scientific Discovery, SD£©¡£¡£¡£¡£¡£¡£¡£ÔÚÎÒ¿´À´£¬£¬£¬£¬£¬³ýÁË¿ÆÑ§ÖÇÄÜ£¨AI for Science, AI4S£©ËùÔÊÐíµÄÖÎÓú°©Ö¢µÈÖî¶àÒæ´¦Ö®Í⣬£¬£¬£¬£¬¿ÆÑ§·¢Ã÷¸üÊÇÍÆÀíÖÇÄܵÄ×îÖÕÄ¥Á·£¬£¬£¬£¬£¬Òò´ËÒ²ÊÇ AI ̽Ë÷µÄ¾ø¶ÔÇ°ÑØ¡£¡£¡£¡£¡£¡£¡£¿£¿£¿£¿£¿£¿£¿ÆÑ§·¢Ã÷ÊÇÒÑÖªÓëδ֪֮¼äÖØ´óµÄÏ໥×÷Ó㬣¬£¬£¬£¬º­¸ÇÁË´Ó¼ÙÉèÌìÉú¡¢ÊµÑéÑéÖ¤µ½ÀíÂÛ×ܽáµÄÈ«Àú³Ì¡£¡£¡£¡£¡£¡£¡£Æä¶Ô AI Ìá³öÁËÈýÖØ¼«ÏÞÌôÕ½£º

ÒÑÖªµÄδ֪£ºµä·¶µÄÈç×éºÏ±¬Õ¨£¬£¬£¬£¬£¬ºÃ±È·Ö×ÓÉè¼Æ»òÖÊÁÏ¿ÆÑ§µÄËÑË÷¿Õ¼ä¸ß´ï 10^60 Á¿¼¶£¬£¬£¬£¬£¬Ô¶³¬¹Å°å±éÀúÄÜÁ¦ £» £»£»£»£»£»Î´ÖªµÄδ֪£º¿ÆÑ§Ì½Ë÷ʵÖÊÉÏÊǶÔÂþÑÜÍ⣨OOD£©ÖªÊ¶µÄ·º»¯£¬£¬£¬£¬£¬ÊǶÔÄ£×Ó´´Á¢Á¦µÄÕæÕýÄ¥Á· £» £»£»£»£»£»Ï£º±ÓëÑÓ³Ù½±Àø£º¿ÆÑ§ÊµÑéµÄÖÜÆÚ³¤¡¢·´ÏìÂý£¬£¬£¬£¬£¬ÊǶÔÇ¿»¯Ñ§Ï°Ëã·¨µÄÑÏËà²âÊԢܡ£¡£¡£¡£¡£¡£¡£

Òò´Ë£¬£¬£¬£¬£¬¿ÆÑ§·¢Ã÷²»µ«ÊÇ AI µÄ×î¼ÑÓ¦Óó¡¾°£¬£¬£¬£¬£¬¸üÊÇÇý¶¯ ¡°Í¨×¨Èںϡ± ÂõÏò AGI µÄ»ù´¡¶¯Á¦¡£¡£¡£¡£¡£¡£¡£

½ÓÏÂÀ´£¬£¬£¬£¬£¬ÎÒÏë·ÖÏíÎÒÃÇΪӦ¶ÔÕâÒ»ÌôÕ½Ìá³öµÄÊÖÒռܹ¹ ¡ª¡ª¡°ÖÇÕß¡±SAGE¡£¡£¡£¡£¡£¡£¡£

ÊÖÒռܹ¹£ºµÝ¹éÑ­»·µÄͨÓÃר¼ÒЭͬ¼Ü¹¹¡°ÖÇÕß¡±SAGE

Ϊ½« ¡°Í¨×¨Èںϡ± Õ½ÂÔת»¯Îª¿ÉÂ䵨µÄÊÖÒռƻ®£¬£¬£¬£¬£¬ÉϺ£ AI ʵÑéÊÒÔÚ 2024 ÄêÌá³öÁË¡°ÖÇÕß¡±SAGE ¼Ü¹¹¡ª¡ª Æä²¢·ÇÈô¸ÉÄ£×ӵļòÆÓ¶ÑÆö£¬£¬£¬£¬£¬¶øÊÇÒ»¸öÖ¼ÔÚÃֺϹãѰ³£»¯ÓëÉî¶Èר¾«ºè¹µµÄͳһÈÏÖªÉú̬ϵͳ¢Ý¡£¡£¡£¡£¡£¡£¡£¸Ã¼Ü¹¹ÓÉÈý¸öÂß¼­ñîºÏµÄÌõÀí×é³É£º

µ×²¿µÄ»ù´¡Ä£×Ó²ãÖÂÁ¦ÓڽṹÉϵÄÖØ¹¹£¬£¬£¬£¬£¬Í¨¹ý½«ÖªÊ¶´¢±¸ÓëÍÆÀíÄÜÁ¦½âñ£¬£¬£¬£¬Îª¸ß½×Òò¹ûÍÆÀíÌṩ¸üÎÞаµÄ ¡°»­²¼¡± £» £»£»£»£»£»ÖÐÐĵÄÈÚºÏЭͬ²ãͨ¹ý÷缯Àú³Ì½±Àø»úÖÆ£¬£¬£¬£¬£¬¶¯Ì¬Ð­µ÷Ö±¾õʽ ¡°¿ì˼Ë÷¡± ÓëÂß¼­ÐÔ ¡°Âý˼Ë÷¡±£¬£¬£¬£¬£¬¾«×¼°Ñ¿Ø·º»¯Óëר¾«µÄ½Ú×à £» £»£»£»£»£»¶¥²ãµÄ̽Ë÷½ø»¯²ãÔò¸¶Óë AI ×Ô¶¯Äܶ¯ÐÔ£¬£¬£¬£¬£¬Íê³É´Ó±»¶¯Êý¾ÝÄâºÏµ½×Ô¶¯ÇéÐÎ̽Ë÷µÄ·¶Ê½×ª±ä¡£¡£¡£¡£¡£¡£¡£

ÖÁ¹ØÖ÷ÒªµÄÊÇ£¬£¬£¬£¬£¬SAGE ¾ø·Ç¾²Ì¬µÄ¼Ü¹¹£¬£¬£¬£¬£¬¶øÊÇÒ»¸öµÝ¹éÔËÐеĻîÌåÉú̬¡£¡£¡£¡£¡£¡£¡£Ëüͨ¹ýË«ÏòÑ­»·ÊµÏÖȫջ½ø»¯£ºÒ»·½Ã棬£¬£¬£¬£¬µ×²ã½âñîµÄ±íÕ÷×Ô϶øÉϵØÖ§³ÖÍÆÀíÕ½ÂÔµÄÌìÉú £» £»£»£»£»£»ÁíÒ»·½Ã棬£¬£¬£¬£¬¶¥²ã×Ô¶¯·¢Ã÷»ñµÃµÄ¸ßË®ÕÑÑ©À¡×ÔÉ϶øÏ»ØÁ÷£¬£¬£¬£¬£¬½«Ì½Ë÷ÖÐµÄ ¡°Î´Öª¡± ת»¯ÎªÐµÄѵÁ·Ðźš£¡£¡£¡£¡£¡£¡£ÕâÖÖ±Õ»·»úÖÆÈ·±£ÁË SAGE ²»µ«ÄÜʵÏÖÄ£×Ó²ÎÊýµÄÓÅ»¯£¬£¬£¬£¬£¬¸üÄÜÍÆ¶¯ÈÏÖªÕ½ÂÔ×Ô¼ºµÄÒ»Á¬½ø»¯¡£¡£¡£¡£¡£¡£¡£

µÝ¹éÑ­»·µÄͨרÈÚºÏÊÖÒռܹ¹¡°ÖÇÕß¡±£¨SAGE£©

»ù´¡Ä£×Ӳ㣺֪ʶÓëÍÆÀíµÄ½â¹¹Ó붯̬ñîºÏ

SAGE µÄµ×²ãÖÂÁ¦ÓÚ½â¾öÏÖÓÐ LLM ½« ¡°ÊÂʵӰÏó¡± Óë ¡°Âß¼­ÍÆÀí¡± »ìÏýµÄÎÊÌâ¡£¡£¡£¡£¡£¡£¡£ÒÔÓ°Ïó½âÂëÆ÷£¨Memory Decoder£©¢ÞΪÀý£¬£¬£¬£¬£¬ËüÕë¶ÔÐԵؽâ¾öÁËÏÖÓдóÄ£×Ӽܹ¹µÄÁ½´óÍç¼²£ºÒ»ÊǼìË÷ÔöÇ¿ÌìÉú£¨RAG£©ÔÚ³¤Îı¾Óï¾³ÍÆÀíÖб£´æµÄÏÔÖøÑÓ³ÙÓë¸ß°º¹¤³Ì±¾Ç® £» £»£»£»£»£»¶þÊÇÁìÓò×Ô˳Ӧȫ²ÎÊý΢µ÷Ëù´øÀ´µÄËãÁ¦ÏûºÄ¼°ÔÖÄÑÐÔÒÅÍüΣº¦¡£¡£¡£¡£¡£¡£¡£

×÷ΪһÖÖԤѵÁ·¡¢¼´²å¼´ÓõÄ×ÔÁ¦×é¼þ£¬£¬£¬£¬£¬Ó°Ïó½âÂëÆ÷Á¢ÒìÐԵؽÓÄÉÓë»ù´¡Ä£×Ó²¢ÐÐÔËÐв¢ÈÚºÏÊä³öÂþÑܵĻúÖÆ¡£¡£¡£¡£¡£¡£¡£ËüÊ×´ÎÓýô´ÕµÄ²ÎÊý»¯Ä£×ÓÌæ»»Á˹Űå·Ç²ÎÊý¼ìË÷Æ÷£¬£¬£¬£¬£¬ÔÚÎÞÐèÐ޸Ļù´¡Ä£×Ó²ÎÊý¡¢ÎÞÔÚÏß¼ìË÷¿ªÏúµÄÌõ¼þÏ£¬£¬£¬£¬£¬ÊµÏÖÁ˸ßЧµÄ֪ʶעÈë¡£¡£¡£¡£¡£¡£¡£ÊµÑéÊý¾ÝÏÔʾ£¬£¬£¬£¬£¬ÆäÍÆÀí¿ªÏú½öΪ»ù´¡Ä£× 1.28 ±¶£¬£¬£¬£¬£¬ÏÔÖøµÍÓÚÏÖÓÐÖ÷Á÷¼Æ»®¡£¡£¡£¡£¡£¡£¡£ÕâÒ»Éè¼ÆÀÖ³ÉÌî²¹ÁË ¡°¸ßÃܶÈ֪ʶ¹©Ó¦¡± Óë ¡°ÍÆÀíÒýÇæ½âñ Ö®¼äµÄÊÖÒպ蹵£¬£¬£¬£¬£¬ÔÚ SAGE ¿ò¼ÜÖÐʵÏÖÁËÍÆÀíÄÜÁ¦Óëºã¾ÃÓ°ÏóµÄ ¡°½âñ¿É¼¯³ÉµÄÍÆÀíÓë֪ʶ¡±£¬£¬£¬£¬£¬Í¬Ê±Ç¿»¯ÁË ¡°ºã¾ÃÓ°Ïó¡± ÄÜÁ¦¡£¡£¡£¡£¡£¡£¡£

Ó°Ïó½âÂëÆ÷£ºÃæÏò´óÓïÑÔÄ£×ÓµÄԤѵÁ·¡¢¼´²å¼´ÓÃÓ°ÏóÌå

Ç¿»¯Ñ§Ï°£ºÅþÁ¬»ù´¡²ãÓë½ø»¯²ãµÄŦ´ø

Ç¿»¯Ñ§Ï°£¨RL£©ÊÇÅþÁ¬ SAGE »ù´¡²ãÓëÈںϲ㡢½ø»¯²ãµÄŦ´ø£¬£¬£¬£¬£¬Ò²ÊÇʵÏÖ ¡°Í¨×¨Èںϡ± µÄ½¹µã¶¯Á¦Ö®Ò»¡£¡£¡£¡£¡£¡£¡ £» £»£»£»£»£»ØÊׯäÑݽøÀú³Ì£¬£¬£¬£¬£¬RL ÂÄÀúÁË´ÓÔçÆÚ¹Ø±ÕÇéÐÎϵIJ©ÞÄ£¨Èç AlphaGo£©£¬£¬£¬£¬£¬ÑݽøÖÁͨ¹ý RLHF ʵÏÖÈËÀàÆ«ºÃ¶ÔÆë£¬£¬£¬£¬£¬ÏÖÔÚÕý´¦ÓÚÒÔ o1 ºÍ DeepSeek-R1 Ϊ´ú±íµÄ¿ÉÑéÖ¤ÍÆÀí£¨RLVR£©½×¶Î£¬£¬£¬£¬£¬²¢ÖÕ½«ÂõÏòÃæÏòÎïÀíÌìÏÂÓë¿ÆÑ§·¢Ã÷µÄ¿ª·ÅʽÌåÑéѧϰмÍÔª¡£¡£¡£¡£¡£¡£¡£

ÊÊÓÃÓÚ¿ÉͨרÈںϵÄÇ¿»¯Ñ§Ï°¼°ÆäÈý´óÖ§Öù

ÔÚ΢¹Û»úÖÆÉÏ£¬£¬£¬£¬£¬RL ±»¹éÄÉΪÈý´óÖ§Öù£º½±ÀøÉè¼Æ×÷Ϊ ¡°Ö¸ÄÏÕ롱£¬£¬£¬£¬£¬Í¨¹ýÏ£º±»ò÷缯ÐźŽ綨ģ×Óר¾«µÄÄ¿µÄ £» £»£»£»£»£»Õ½ÂÔÓÅ»¯×÷Ϊ ¡°ÒýÇæ¡±£¬£¬£¬£¬£¬º­¸Ç´Ó PPO µ½ GRPO µÄËã·¨µü´ú£¬£¬£¬£¬£¬Çý¶¯Ä£×Ó¸ßЧ¸üР£» £»£»£»£»£»²ÉÑùÓë̽Ë÷Ôò¾öÒéÁËÄ£×ÓÔÚÖØ´óËÑË÷¿Õ¼äÖеĵ¼º½Â·¾¶¢ß¡£¡£¡£¡£¡£¡£¡£

¼øÓÚ²î±ðʹÃü¶Ô RL ÉèÖõÄÐèÇó¸÷Ò죬£¬£¬£¬£¬¹¹½¨ÏµÍ³µÄ½¹µãÊÖÒÕÌôÕ½ÔÚÓÚͳһ£ºÎÒÃÇÔõÑù½«¶àÑùÐÔµÄ×î¼ÑµÄ½±Àø»úÖÆ¡¢Õ½ÂÔÓÅ»¯Óë²ÉÑù̽Ë÷ÕûºÏΪһ¸öЭµ÷Ò»ÖµÄϵͳ£¬£¬£¬£¬£¬´Ó¶ø´òÔì³öÕæÕýµÄ ¡°¿ÉÉî¶Èרҵ»¯Í¨ÓÃÄ£×Ó¡±£¿£¿£¿£¿£¿£¿£¿

ÈÚºÏЭͬ²ã£ºÇ¿»¯Ñ§Ï°Çý¶¯µÄÉî¶ÈÍÆÀí½ø»¯

ÔÚ SAGE ¼Ü¹¹ÖУ¬£¬£¬£¬£¬ÈÚºÏЭͬ²ã³ÐÔØ×ÅЭµ÷ ¡°Ö±¾õ¿ì˼Ë÷¡± Óë ¡°Âß¼­Âý˼Ë÷¡± µÄ½¹µãÖ°ÄÜ£¬£¬£¬£¬£¬¶øÇ¿»¯Ñ§Ï°£¨RL£©ÔòÊÇʵÏÖÕâÒ»¶¯Ì¬Ð­Í¬µÄÒªº¦ÇÅÁº¡£¡£¡£¡£¡£¡£¡£ÎªÁ˹¹½¨Ò»¸öÕæÕýµÄ ¡°¿ÉÉî¶Èרҵ»¯Í¨ÓÃÄ£×Ó¡±£¬£¬£¬£¬£¬±ØÐèսʤ¹Å°å RL ÔÚÖØ´óÍÆÀíʹÃüÖÐÃæÁÙµÄÈý´ó½¹µãÌôÕ½£º¸ß°ºµÄ¼àÊÓ±¾Ç®¡¢ÑµÁ·Àú³ÌÖеÄìØÌ®ËõÒÔ¼°¼òµ¥Æð¾¶µÄģʽÍ߽⡣¡£¡£¡£¡£¡£¡£Îª´Ë£¬£¬£¬£¬£¬ÎÒÃÇÔڸòãÒýÈëÁËÈýÏî¾ßÓз¶Ê½ÒâÒåµÄËã·¨Á¢Ò죬£¬£¬£¬£¬Ö¼ÔÚ¹¹½¨÷缯µÄ½±Àø»úÖÆ¡¢Î¬³ÖÒ»Á¬µÄ̽Ë÷ÄÜÁ¦ÒÔ¼°Òý·¢ÍÆÀí·¾¶µÄ¶àÑùÐÔ¡£¡£¡£¡£¡£¡£¡£

Òþʽ½±ÀøÇ¿»¯Ñ§Ï°Ëã·¨£¨PRIME£©£ºÍ»ÆÆ¸ßÃܶȼàÊӵı¾Ç®ã£ÂÛ

¸ß¶Èר¼Ò»¯µÄÄ£×ÓÓëÈËÀàר¼ÒÔÚѧϰ»úÖÆÉϾßÓÐÏàËÆÐÔ£º×¨¼Ò»¯Ä£×ÓÔÚѵÁ·Àú³ÌÖÐÐèÒª¸ü÷缯µÄ·´ÏìÐÅÏ¢¡£¡£¡£¡£¡£¡£¡£¹ØÓÚ ¡°Í¨×¨Èںϡ± ´óÄ£×Ó¶øÑÔ£¬£¬£¬£¬£¬Òª½â¾ö¿ÆÑ§·¢Ã÷Öеij¤Á´ÌõÍÆÀíÎÊÌ⣬£¬£¬£¬£¬½öÒÀÀµ×îÖÕЧ¹ûµÄÏ£º±½±ÀøÍùÍù×óÖ§ÓÒç©£¬£¬£¬£¬£¬Ä£×Ó¼±Ðè÷缯µÄÖð²½¼àÊÓÐźš£¡£¡£¡£¡£¡£¡£È»¶ø£¬£¬£¬£¬£¬¹Å°åµÄ½â¾ö¼Æ»®ÒÀÀµÓÚÀú³Ì½±ÀøÄ£×Ó£¨PRM£©£¬£¬£¬£¬£¬ÕâÒªÇó¶Ôº£Á¿ÍÆÀí°ì·¨¾ÙÐÐÈ˹¤Ï¸Á£¶È±ê×¢£¬£¬£¬£¬£¬Æä±¾Ç®Ö®¸ß°º£¬£¬£¬£¬£¬Ê¹µÃ¹æÄ £» £»£»£»£»£»¯À©Õ¹ÏÕЩ³ÉΪ²»¿ÉÄÜ¡£¡£¡£¡£¡£¡£¡£

Õë¶ÔÕâÒ» ¡°¸ßÃܶȼàÊÓÐèÇó¡± Óë ¡°¸ß°º±ê×¢±¾Ç®¡± Ö®¼äµÄì¶Ü£¬£¬£¬£¬£¬ÎÒÃÇÌá³öÁË PRIME Ëã·¨¢à £¬£¬£¬£¬£¬Ö¼ÔÚ´ÓÀíÂÛ²ãÃæÍÆµ¼²¢»ñÈ¡ ¡°Ãâ·Ñ¡± µÄÀú³Ì½±Àø¡£¡£¡£¡£¡£¡£¡£Æä½¹µã¶´²ìÔÚÓÚ£¬£¬£¬£¬£¬Ê¹ÓÃÕ½ÂÔÄ£×ÓÓë²Î¿¼Ä£×ÓÖ®¼äµÄͳ¼Æ²î±ð¡£¡£¡£¡£¡£¡£¡£Í¨¹ý½«Ä£×ÓѵÁ·Ä¿µÄÉ趨Ϊ»ùÓÚÁ½Õß¶ÔÊýËÆÈ»±ÈµÄЧ¹û½±ÀøÄ£×Ó£¬£¬£¬£¬£¬ÎÒÃÇ´ÓÊýѧ·½ÃæÖ¤Êµ£¬£¬£¬£¬£¬¸ÃÄ£×ÓÄܹ»ÒþʽµØÏ°µÃ Q º¯Êý¡£¡£¡£¡£¡£¡£¡£ÕâÒâζ×Å£¬£¬£¬£¬£¬ÖÇÄÜÌåÔÚÎÞÐèÏÔʽѵÁ·ÖØ´óµÄ PRM Ä£×ÓµÄÇéÐÎÏ£¬£¬£¬£¬£¬¼´¿ÉÔÚÍÆÀíµÄÿһ¸ö°ì·¨ÖУ¬£¬£¬£¬£¬Í¨¹ýÅÌËãÐж¯ÔÚÄ¿½ñ״̬ϵÄÓÅÁÓ£¬£¬£¬£¬£¬Ö±½ÓÍÆµ¼³ö÷缯µÄ¡¢Ö𲽵Ľ±ÀøÐźš£¡£¡£¡£¡£¡£¡£

Òþʽ½±ÀøÇ¿»¯Ñ§Ï°Ëã·¨£¨PRIME£©

ÕâÒ»Á¢Òì´øÀ´Á˶àά¶ÈµÄÏÔÖøÓÅÊÆ£º

ÅÌËãЧÂʵı¼ÌÚ£ºÓë Math-Shepherd µÈÒÀÀµ×ÔÁ¦ PRM Ä£×ÓµÄÒªÁìÏà±È£¬£¬£¬£¬£¬PRIME ÔÚÍÆÀí½×¶ÎÎÞÐèÌØÁíÍâÄ£×ÓŲÓÿªÏú£¬£¬£¬£¬£¬Ö±½ÓʹÓÃÌìÉúÄ£×Ó×Ô¼ºµÄ¸ÅÂÊÂþÑܼ´¿É»ñµÃ·´Ï죬£¬£¬£¬£¬¼«´óµØÌáÉýÁËÅÌËãЧÂÊ £» £»£»£»£»£»ÏµÍ³¼Ü¹¹µÄ¿ÉÀ©Õ¹ÐÔ£ºÔÚ SAGE µÄϵͳʵÏÖÖУ¬£¬£¬£¬£¬PRIME ¼Æ»®Õ¹ÏÖ³ö¼«Ç¿µÄ¹¤³ÌÈÍÐÔ¡£¡£¡£¡£¡£¡£¡£ÎÒÃǽ«Õ½ÂÔÄ£×ÓÓëÒþʽ PRM ¾ÙÐÐÁª¶¯£¬£¬£¬£¬£¬ÒÀÍÐЧ¹ûÑéÖ¤Æ÷ºÍǰÐò°ì·¨²ú³öµÄ×ÔÓÉÀú³Ì½±Àø£¬£¬£¬£¬£¬¹¹½¨Á˸ßЧµÄÔÚÏ߸üбջ· £» £»£»£»£»£»¼«ÖµÄÊý¾ÝЧÂÊ£ºÊµÑéÅú×¢£¬£¬£¬£¬£¬PRIME ¼Æ»®½öÐè SOTA Ä£×Ó 1/10 µÄѵÁ·Êý¾ÝÁ¿£¬£¬£¬£¬£¬¼´¿ÉµÖ´ïÏ൱µÄÐÔÄÜˮƽ£¬£¬£¬£¬£¬¼«´óµØ½µµÍÁ˶ԸßÖÊÁ¿±ê×¢Êý¾ÝµÄÒÀÀµ¡£¡£¡£¡£¡£¡£¡£

»ù×¼²âÊÔЧ¹ûÓÐÁ¦µØÑéÖ¤ÁË PRIME µÄÓÐÓÃÐÔ£ºÔÚ AIME 2024 Êý¾Ý¼¯ÉÏ£¬£¬£¬£¬£¬Ä£×Ó׼ȷÂÊÌáÉýÁË 23.4% £» £»£»£»£»£»ÔÚ AMC Êý¾Ý¼¯ÉÏÌáÉýÁË 27.7% £» £»£»£»£»£»ÔÚ MATH-500 µÈȨÍþ²âÊÔÖÐҲȡµÃÁËÏÔÖøÔöÌí¡£¡£¡£¡£¡£¡£¡£ÕâһϵÁÐÊý¾Ý³ä·Ö֤ʵ£¬£¬£¬£¬£¬Í¨¹ýÒþʽ»úÖÆ¹¹½¨µÄŨÃܽ±Àø£¬£¬£¬£¬£¬Äܹ»ÓÐÓÃÇý¶¯Ä£×ÓÍ»ÆÆÖØ´óÍÆÀíµÄÆ¿¾±¡£¡£¡£¡£¡£¡£¡£

Ç¿»¯Ñ§Ï°µÄìØ»úÖÆ£º×èÖ¹ ¡°Ì«¹ý×ÔÐÅ¡± µ¼ÖÂ̽Ë÷Ö¹²½

ר¼Ò»¯Ä£×ÓµÄѵÁ·²»µ«ÐèÒª·´Ï죬£¬£¬£¬£¬¸üÐèÒªÒ»Á¬Ò»Ö±µÄѧϰ¡£¡£¡£¡£¡£¡£¡£ÔÚÉîÈëÑо¿ÓÃÓÚÍÆÀíµÄÇ¿»¯Ñ§Ï°Ê±£¬£¬£¬£¬£¬ÎÒÃÇÕ¹ÏÖÁËÒ»¸ö×è°­Ä£×Ó½ø»¯µÄ¸ùÌìÐÔÕϰ­ ¡ª¡ªìØÌ®Ëõ¡£¡£¡£¡£¡£¡£¡£Í¨Ë׵ؽ²£¬£¬£¬£¬£¬ÕâµÈͬÓÚ½â¾öÔõÑùÈÃͨÓÃÄ£×ÓÔÚר¼Ò»¯µÄÀú³ÌÖУ¬£¬£¬£¬£¬Ê¼ÖÕ¼á³Ö̽Ë÷ÓëºÃÆæÐÄ£¬£¬£¬£¬£¬ÈÃÄ£×ӺͶ¥¼¶ÈËÀàר¼ÒÒ»ÑùÔÚרҵÎÊÌâµÄÌôÕ½ÉÏ×èÖ¹¹ýÔçÌ«¹ý×ÔÐÅ£¬£¬£¬£¬£¬¶øÊÇ ¡°stay hungry, stay foolish¡±£¨ÇóÖªÈô¼¢£¬£¬£¬£¬£¬ÐéÐÄÈôÓÞ£©¡£¡£¡£¡£¡£¡£¡£

ÔÚѵÁ·Àú³ÌÖУ¬£¬£¬£¬£¬Ëæ×ÅÄ£×ÓÐÔÄܵįðÔ´ÌáÉý£¬£¬£¬£¬£¬Õ½ÂÔìØÍùÍù»á¼±¾çϽµ¡£¡£¡£¡£¡£¡£¡£ÕâÖÖϽµÒâζ×ÅÄ£×Ó¶ÔÆäÊä³öµÄÖÃÐŶȿìËÙÌá¸ß£¬£¬£¬£¬£¬µ¼ÖÂÆä¹ýÔçµØÊÕÁ²ÓÚ¾Ö²¿×îӎ⣬£¬£¬£¬£¬´Ó¶øËðʧÁË̽Ë÷¸üÓÅÍÆÀí·¾¶µÄ¿ÉÄÜÐÔ¡£¡£¡£¡£¡£¡£¡£ÊµÑéÊý¾ÝÏÔʾ£¬£¬£¬£¬£¬ìصÄÏûºÄÖ÷Òª¼¯ÖÐÔÚѵÁ·µÄǰÊý°Ù²½£¬£¬£¬£¬£¬ÒÔºóÄ£×ÓµÄÐÔÄÜÌáÉý±ãѸËÙ½øÈë±ß¼ÊÐ§ÒæµÝ¼õ½×¶Î¡£¡£¡£¡£¡£¡£¡£ÕâÖÖÕ÷Ïó¼«ËÆÈËÀàÈÏÖªÖÐµÄ ¡°Ì«¹ý×ÔÐÅ¡±£¬£¬£¬£¬£¬¼´Òò×ÔÂú¶ø×èÖ¹Á˶ÔÎÊÌâϸ΢²î±ðµÄ×Ô¶¯Ì½Ë÷ ¡ª¡ª ¶øÕâÖÖ×Ô¶¯Ì½Ë÷£¬£¬£¬£¬£¬Ç¡Ç¡ÊÇͨÓÃÄ£×Ó½ø»¯ÎªÄܲ¶»ñÉî²ã¼ÍÂÉµÄ ¡°×¨¾«Ä£×Ó¡± µÄÒªº¦ËùÔÚ¡£¡£¡£¡£¡£¡£¡£

ΪÏàʶ¾öÕâÒ»ÎÊÌ⣬£¬£¬£¬£¬ÎÒÃÇÉîÈë̽ÌÖÁËìØÓë½±ÀøÖ®¼äµÄȨºâ»úÖÆ£¬£¬£¬£¬£¬²¢·¢Ã÷ÁËÒ»¸öÒªº¦µÄ¶¨Á¿¹ØÏµ£ºÑéÖ¤ÐÔÄÜ£¨R£©ÓëìØ£¨H£©·ºÆðÏÔÖøµÄ¶ÔÊýÏßÐÔÏà¹Ø¢á¡£¡£¡£¡£¡£¡£¡£ÕâÒ»¾«Á·¶øÉî¿ÌµÄ½áÂÛΪѵÁ·¼Æ»®µÄÓÅ»¯Ö¸Ã÷ÎúÆ«Ïò£º¹¹½¨¿ÉÀ©Õ¹ÍÆÀí RL ¿ò¼ÜµÄÄѵ㣬£¬£¬£¬£¬²»ÔÚÓÚ´¿´â¶ÑÆöѵÁ·Ê±³¤£¬£¬£¬£¬£¬¶øÔÚÓÚ¶ÔìØÏûºÄµÄϸÄ廯ÖÎÀí£¬£¬£¬£¬£¬È·±£Ä£×ÓÔÚѵÁ·È«ÖÜÆÚÄÚ±£´æ×ã¹»µÄ²»È·¶¨ÐÔ£¬£¬£¬£¬£¬ÒÔÇý¶¯Ò»Á¬µÄ̽Ë÷¡£¡£¡£¡£¡£¡£¡£

ÎÒÃÇÌá³öÁËÒ»ÖÖ¾«×¼»¯¡¢¾Ö²¿»¯ÇÒÇáÁ¿»¯µÄìØ¿ØÖƼƻ®£ºÕë¶ÔÕâÀà±ê¼Ç¿ªÕ¹Ñ¡ÔñÐÔµ÷¿Ø£¨Èç½ÓÄÉ Clip-Cov¡¢KL-Cov µÈÒªÁ죩£¬£¬£¬£¬£¬Äܹ»¸æ¿¢¾Ö²¿¡¢ÇáÁ¿µÄìØ¿ØÖÆÐ§¹û£¬£¬£¬£¬£¬¼È°ü¹ÜÄ£×Ó̽Ë÷ÐÔ²»ÊÜË𣬣¬£¬£¬£¬ÓÖ²»»á×ÌÈÅÕý³£ÓÅ»¯Á÷³Ì¡£¡£¡£¡£¡£¡£¡£¸ÃÒªÁìʵÏÖÁ˶ÔìØµÄ¾Ö²¿¿ØÖÆ£¬£¬£¬£¬£¬¼È°ü¹ÜÁËÄ£×ÓµÄ̽Ë÷ÐÔ²»ÊÜË𣬣¬£¬£¬£¬ÓÖ×èÖ¹Á˶ÔÕý³£ÓÅ»¯Á÷³ÌµÄ×ÌÈÅ¡£¡£¡£¡£¡£¡£¡£Ó¦ÓøÃÕ½ÂԺ󣬣¬£¬£¬£¬Ä£×ÓÔÚ¼á³Ö¸ß̽Ë÷ÄÜÁ¦µÄͬʱ£¬£¬£¬£¬£¬ÏÔÖøÌáÉýÁËÏÂÓÎʹÃüµÄ׼ȷÂÊ¡£¡£¡£¡£¡£¡£¡£ÕâÒ»ÒªÁìÒѱ»ÊµÑéÊҵġ°ÊéÉú¡±¿ÆÑ§¶àģ̬´óÄ£×Ó Intern-S1 µÈ¶à¸öÍ·²¿»ú¹¹½ÓÄÉÓ¦Ó㬣¬£¬£¬£¬ÆäÏà¹ØÐ§¹û¸üÓÉ˹̹¸£ Yejin Choi ½ÌÊÚÔÚ 2025 ÄêÉñ¾­ÐÅÏ¢´¦Öóͷ£ÏµÍ³´ó»á£¨NeurIPS£©ÉϾÙÐÐÁËÖØµãÐðÊö¡£¡£¡£¡£¡£¡£¡£

Ç¿»¯Ñ§Ï°µÄìØ»úÖÆ

Æ¥Åä´óÓïÑÔÄ£×ÓÍÆÀíµÄ½±ÀøÂþÑÜ£¨FlowRL£©£ºÊµÏÖר¼Ò»¯Ä£×ÓÄÜÁ¦¶àÔª»¯

ÕæÕýµÄר¼Ò²»µ«Äܽâ¾öÎÊÌ⣬£¬£¬£¬£¬¸üÄÜÄÜΪͳһ¸öÎÊÌâÌṩ¶àÖÖ½â¾ö¼Æ»®£¬£¬£¬£¬£¬×¨¼Ò»¯Ä£×ÓÒàÊÇÔÆÔÆ¡£¡£¡£¡£¡£¡£¡£È»¶ø£¬£¬£¬£¬£¬ÏÖÓеıê׼ǿ»¯Ñ§Ï°ÒªÁ죨Èç PPO¡¢GRPO£©ÆÕ±éÒÔ ¡°½±Àø×î´ó»¯¡± Ϊ¼òµ¥Ä¿µÄ¡£¡£¡£¡£¡£¡£¡£ÕâÖÖµ¼ÏòÔÚÖØ´óÍÆÀíʹÃüÖм«Ò×µ¼ÖÂģʽÍ߽⣬£¬£¬£¬£¬¼´Ä£×ÓÇãÏòÓÚÖØ¸´ÊÕÁ²ÖÁ¼òµ¥µÄ¡¢ÒÑÖªµÄÀֳɷ¾¶£¬£¬£¬£¬£¬¶øºöÂÔÁËÆäËûDZÔڵĸüÓŽâ»ò¶àÑù»¯½â·¨¡£¡£¡£¡£¡£¡£¡£

¹Å°å RL ÒªÁìÌìÉúµÄÂþÑÜÓëÄ¿µÄÂþÑÜÖ®¼äµÄ KL É¢¶È¸ß´ï 8.68£¬£¬£¬£¬£¬ÌåÏÖΪ¼«¶ËµÄ¼â·å£¬£¬£¬£¬£¬Òâζ×ÅÄ£×Ó̽Ë÷¿Õ¼äµÄ¼«¶ËÏÁÕ­¡£¡£¡£¡£¡£¡£¡£ÎªÁ˸¶ÓëÄ£×ÓÕæÕýµÄר¼Ò¼¶Í·ÄÔ¶àÑùÐÔ£¬£¬£¬£¬£¬ÎÒÃÇÔÚÈںϲãÒýÈëÁËFlowRL¢â£¬£¬£¬£¬£¬ÕâÊÇÒ»Ïî½è¼øÌìÉúÁ÷ÍøÂ磨GFlowNets£©Í·ÄÔµÄÁ¢ÒìÊÂÇ飬£¬£¬£¬£¬±ê¼Ç×ÅÇ¿»¯Ñ§Ï°ÓÅ»¯Âß¼­µÄ·¶Ê½×ª±ä¡£¡£¡£¡£¡£¡£¡£

FlowRL µÄ½¹µãÔÚÓÚ½«Ñ§Ï°Ä¿µÄ´Ó ¡°½±Àø×î´ó»¯¡± ÖØ¹¹Îª ¡°ÂþÑÜÆ¥Å䡱¡£¡£¡£¡£¡£¡£¡£Ä£×Ó²»ÔÙ½ö½ö×·Öð¼òµ¥µÄ¸ß·ÖÃÕµ×£¬£¬£¬£¬£¬¶øÊÇÖÂÁ¦ÓÚѧϰËùÓÐÓÐÓÃÍÆÀí·¾¶µÄ¸ÅÂÊÂþÑÜ¡£¡£¡£¡£¡£¡£¡£

ÂþÑÜÄâºÏ£ºFlowRL ÌìÉúµÄÂþÑÜÄܹ»²¶»ñÄ¿µÄÂþÑÜÖеľø´ó´ó¶¼¸ÅÂÊÖÊÁ¿£¬£¬£¬£¬£¬ÄâºÏ¶à¸öģ̬¡£¡£¡£¡£¡£¡£¡£Èç×ó²àƽ»¬ÇúÏßËùʾ£¬£¬£¬£¬£¬Æä KL É¢¶È´ó·ù½µµÍÖÁ 0.11£¬£¬£¬£¬£¬ÏÔÖøÓÅÓڹŰåÒªÁì £» £»£»£»£»£»¶àÑùÐÔÌìÉú£ºÏ°µÃµÄÕ½ÂÔÔÚÍÆÀíÀú³ÌÖÐÄܹ»×ÔÈ»µØÔö½ø¸ü¶àÑù»¯Â·¾¶µÄÌìÉú£¬£¬£¬£¬£¬´Ó¶øÔÚÃæÁÙ ¡°Î´ÖªµÄδ֪¡± ʱ¾ß±¸¸üÇ¿µÄ³°ôÐÔ¡£¡£¡£¡£¡£¡£¡£

°¸ÀýÏÔʾ£¬£¬£¬£¬£¬ÔÚ´¦Öóͷ£Í³Ò»µÀÊýÑ§ÍÆÀíÌâʱ£¬£¬£¬£¬£¬GRPO Ä£×ÓÏÝÈëÁËÍ·ÄÔËÀÑ­»·£¬£¬£¬£¬£¬ÍÆÀíÀú³ÌÖØ¸´ÇÒ×îÖÕδÄÜÇó½â £» £»£»£»£»£»¶ø FlowRL Ä£×ÓÔòÀÖ³É̽Ë÷Á˶àÑù»¯µÄÍÆÀí·¾¶£¬£¬£¬£¬£¬×îÖյóöÁË׼ȷÃÕµ× 721¡£¡£¡£¡£¡£¡£¡£

ÕûÌåʵÑéЧ¹û½øÒ»²½Ö¤ÊµÁË FlowRL µÄÓÅÔ½ÐÔ£º

׼ȷÂÊÌáÉý£ºÔÚ 32B Ä£×ÓµÄѵÁ·Ìõ¼þÏ£¬£¬£¬£¬£¬FlowRL ÔÚÊýÑ§ÍÆÀíʹÃüÖÐÈ¡µÃÁË 48.39% µÄ׼ȷÂÊ£¬£¬£¬£¬£¬½Ï GRPO ÌáÉý 10 ¸ö°Ù·Öµã£¬£¬£¬£¬£¬½Ï PPO ÌáÉý 5.1 ¸ö°Ù·Öµã £» £»£»£»£»£»¾ºÈü¼¶ÌåÏÖ£º»ùÓÚ´¿¿ªÔ´Êý¾ÝѵÁ·ºó£¬£¬£¬£¬£¬FlowRL ÔÚ CodeForces ƽ̨µÄÆÀ¼¶µÖ´ï 1549 ·Ö£¬£¬£¬£¬£¬ÐÔÄÜÖ±±Æ o1-preview ˮƽ £» £»£»£»£»£»¶àÑùÐÔ±¶Ôö£ºFlowRL ÌìÉúµÄ½â¾ö¼Æ»®¶àÑùÐÔÆÀ·Ö¸ß´ï 2.28£¬£¬£¬£¬£¬Ô¼Îª PPO µÄ 2 ±¶¡£¡£¡£¡£¡£¡£¡£

Æ¥Åä´óÓïÑÔÄ£×ÓÍÆÀíµÄ½±ÀøÂþÑÜ£¨FlowRL£©

̽Ë÷½ø»¯²ã£º´Ó±»¶¯ÄâºÏµ½×Ô¶¯ÈÏ֪̽Ë÷

SAGE ¼Ü¹¹µÄ¶¥²ã̽Ë÷½ø»¯²ã³ÐÔØ×ÅͨÍù AGI ×îÒªº¦µÄÔ¸¾° ¡ª¡ª ´òÔìÒ»¸ö¾ß±¸×ÔÑÝ»¯ÄÜÁ¦µÄ ¡°¿ÉÉî¶Èרҵ»¯Í¨ÓÃÄ£×Ó¡±¡£¡£¡£¡£¡£¡£¡£ÕâÒ»²ãµÄ½¹µãÌôÕ½ÔÚÓÚ£¬£¬£¬£¬£¬ÔõÑùÈÃͨÓÃÄ£×Ó²»µ«ÔÚ¼òµ¥Ê¹ÃüÉÏʵÏÖÉî¶Èר¾«£¬£¬£¬£¬£¬¸üÄÜÔÚ´ó¹æÄ£Ê¹Ãü¼¯ÒÔÖÂÖØ´óµÄÎïÀíÌìÏÂÖУ¬£¬£¬£¬£¬Í¨¹ýÒ»Á¬µÄ½»»¥Óë·´ÏìʵÏÖ×ÔÎÒµü´ú¡£¡£¡£¡£¡£¡£¡£ÎªÁËÓ¦¶ÔÕâÒ»ÌôÕ½£¬£¬£¬£¬£¬ÎÒÃÇ´ÓÐźţ¨Signal£©¡¢¹æÄ££¨Scale£©ÓëÂ䵨£¨Ground£©Èý¸öÒªº¦Î¬¶È³ö·¢£¬£¬£¬£¬£¬¹¹½¨ÁËÒ»Ì×ÍêÕûµÄ½ø»¯»úÖÆ¡£¡£¡£¡£¡£¡£¡£

ÐźÅά¶È£º²âÊÔʱǿ»¯Ñ§Ï°£¨TTRL£©Óë×ÔÎÒ½ø»¯

ÔÚÍÆÀí²âÊԽ׶Σ¬£¬£¬£¬£¬Ä£×ÓÃæÁÙµÄ×î´óÄæ¾³ÔÚÓÚѵÁ·Êý¾ÝÓë²âÊÔÊý¾ÝÖ®¼äµÄÂþÑÜÆ«ÒÆ¡£¡£¡£¡£¡£¡£¡£Ò»µ©Ê§È¥ÕæÊµ±êÇ©µÄÖ¸µ¼£¬£¬£¬£¬£¬¹Å°åÄ£×Ó±ã×èÖ¹ÁËѧϰ³ÌÐò¡£¡£¡£¡£¡£¡£¡£È»¶ø£¬£¬£¬£¬£¬ÕæÕýµÄ ¡°×¨¼Ò¡±¡ª¡ª ÓÌÈçÈËÀàÎïÖÖÒ»Ñù ¡ª¡ª Ó¦µ±¾ß±¸ÔÚÈκÎδ֪¾³¿öÏÂÒ»Á¬Ñ§Ï°Ë³Ó¦µÄÄÜÁ¦¡£¡£¡£¡£¡£¡£¡£

Õë¶ÔÕâһʹµã£¬£¬£¬£¬£¬ÎÒÃÇÌá³öÁ˲âÊÔʱǿ»¯Ñ§Ï°£¨Test-Time Reinforcement Learning, TTRL£©¿ò¼Ü? £¬£¬£¬£¬£¬Æä½¹µã¶´²ì½¨ÉèÔÚÒ»¸ö¾«Á·µÄ¼ÙÉèÖ®ÉÏ£º¹²Ê¶¼´Òâζ×Å׼ȷÐÔ£¨Consensus implies correctness£©¡£¡£¡£¡£¡£¡£¡£

Ïêϸ¶øÑÔ£¬£¬£¬£¬£¬TTRL ÔÚÍÆÀíÀú³ÌÖжԶà¸öºòÑ¡½â¾ö¼Æ»®¾ÙÐвÉÑù£¬£¬£¬£¬£¬²¢½«´ó¶¼Í¶Æ±µÄЧ¹û×÷Ϊ ¡°ÊðÀí½±Àø¡±£¬£¬£¬£¬£¬½ø¶øÊ¹ÓòâÊÔÊý¾ÝÁ÷Ö±½Ó¶ÔÄ£×Ó²ÎÊý¾ÙÐÐÔÚÏ߸üС£¡£¡£¡£¡£¡£¡£ÕâÒ»ÒªÁìÔÚÊÖÒÕʵÏÖÉϾ߱¸¼«ÖµÄÇáÁ¿»¯ÌØÕ÷£¬£¬£¬£¬£¬½öÐè²»µ½ 20 ÐдúÂ룬£¬£¬£¬£¬¼´¿É½«ÈκÎÍÆÀí¹ì¼£×ª»¯ÎªÓÐÓõÄѵÁ·ÐźÅ£¬£¬£¬£¬£¬ÊµÏÖÁËÄ£×ÓÔÚÎÞ¼àÊÓÇéÐÎÏ嵀 ¡°×ÔÎÒ¾ÙÖ¤¡± Óë ¡°×ÔÎÒÔöÇ¿¡±¡£¡£¡£¡£¡£¡£¡£

²âÊÔʱǿ»¯Ñ§Ï°Óë×ÔÎÒ½ø»¯£¨TTRL£©

ʵ²âÊý¾ÝÑéÖ¤ÁË TTRL µÄ¾ªÈËDZÁ¦£º

ÐÔÄÜÔ¾Éý£ºÔÚ AIME 2024 Êý¾Ý¼¯ÉÏ£¬£¬£¬£¬£¬´îÔØ TTRL µÄ Qwen-2.5-Math-7B Ä£×Ó׼ȷÂÊʵÏÖÁË 159% µÄÏà¶ÔÌáÉý £» £»£»£»£»£»×ÔÎÒÓâÔ½£ºTTRL ÓÅ»¯ºóµÄÄ£×ÓÕ¹ÏÖ³öÁË ¡°Çà³öÓÚÀ¶¡± µÄÌØÕ÷£¬£¬£¬£¬£¬ÆäÐÔÄܲ»µ«ÓâÔ½ÁË×ÔÉíµÄ ¡°×îÓÅ N ²ÉÑù¡± »ù×¼Ïߣ¬£¬£¬£¬£¬ÉõÖÁÆÈ½üÁËʹÓôøÕæÊµ±êǩѵÁ·µÄÀíÂÛÉÏÏÞ£¨Oracle »ùÏߣ© £» £»£»£»£»£»Ç¿·º»¯ÐÔ£ºÔÚ AMC¡¢MATH-500 µÈδ¼û¹ýµÄȨÍþ»ù×¼²âÊÔÖУ¬£¬£¬£¬£¬Ä£×ÓͬÑùÌåÏÖ³öÇ¿¾¢µÄ·º»¯ÄÜÁ¦¡£¡£¡£¡£¡£¡£¡£

TTRL µÄÀÖ³É֤ʵÎúÖÇÄÜÌå¾ß±¸×ÔÖ÷ÂÝÐýʽÉÏÉýµÄÉú³¤Ç±Á¦£¬£¬£¬£¬£¬Îª SAGE ¼Ü¹¹ÖеÄ×ÔÎÒ½ø»¯ÌṩÁËÒ»Ìõ¾«Á·¸ßЧµÄ·¾¶¡£¡£¡£¡£¡£¡£¡£

¹æÄ£Î¬¶È£ºInternBootcamp ÓëʹÃüÀ©Õ¹¶¨ÂÉ

ÔÚ½â¾öÁË ¡°Ôõôѧ¡± µÄÐźÅÎÊÌâºó£¬£¬£¬£¬£¬±ØÐè»Ø¸² ¡°ÔÚÄÄѧ¡± µÄ¹æÄ£ÎÊÌâ¡£¡£¡£¡£¡£¡£¡£Í¨×¨ÈÚºÏÄ£×Ó²»µ«ÐèÒªÔÚ¼òµ¥Ê¹ÃüÉÏͨ¹ý ¡°Âý˼Ë÷¡± ʵÏÖר¾«£¬£¬£¬£¬£¬¸üÐèÒªÔڳɰÙÉÏǧ¸öʹÃüÉÏͬʱʵÏÖÄÜÁ¦ÊÊÅä¡£¡£¡£¡£¡£¡£¡£±ðµÄ£¬£¬£¬£¬£¬ÎÒÃÇ»¹Ï£Íû̽Ë÷Ò»¸ö¸üÉî¿ÌµÄÎÊÌ⣺µ±²âÊÔʹÃüµÄÊýÄ¿Óë¶àÑùÐÔͬ²½À©Ôöʱ£¬£¬£¬£¬£¬ÊÇ·ñ±£´æ×¨ÃÅÕë¶ÔÔÚ²âÊÔÇéÐÎÏ¡¢Õë¶ÔʹÃüÊýÄ¿µÄ Scaling Law£¿£¿£¿£¿£¿£¿£¿

Ϊ´Ë£¬£¬£¬£¬£¬ÎÒÃÇÑз¢ÁË´ó¹æÄ£¡¢±ê×¼»¯¡¢¿ÉÀ©Õ¹µÄ½»»¥ÑéÖ¤ÇéÐÎ ¡ª¡ªInternBootcamp?¡£¡£¡£¡£¡£¡£¡£

×÷ΪÊ׸öÁýÕÖ 8 ´óʹÃüÖֱ𡢳¬ 1000 ÖÖ¶àÑù»¯ÇéÐÎµÄÆ½Ì¨£¬£¬£¬£¬£¬InternBootcamp Ö§³ÖÔÚÖ¸¶¨ÇéÐÎÖпªÕ¹´ó¹æÄ£Ç¿»¯Ñ§Ï°ÑµÁ·¡£¡£¡£¡£¡£¡£¡£ÆäÆæÒìµÄ ¡°Ê¹ÃüÓëÑéÖ¤º¯Êý×Ô¶¯ÌìÉú¡± ÄÜÁ¦£¬£¬£¬£¬£¬Ê¹µÃÓû§Äܹ»±ã½ÝµØ½«µç·Éè¼ÆµÈרҵÁìÓòʹÃüת»¯Îª¿ÉÑéÖ¤ÇéÐΣ¬£¬£¬£¬£¬Í¨¹ý·ÂÕæÊÖ¶ÎÍêЧ¹û¹ûºËÑé¡£¡£¡£¡£¡£¡£¡£

InternBootcamp ÁýÕÖ 8 ´óʹÃüÖֱ𡢳¬ 1000 ÖÖ¶àÑù»¯Ê¹ÃüÇéÐÎ

»ùÓÚ InternBootcamp µÄʵÑéÕ¹ÏÖÁËÁ½¸öÖ÷ÒªÕ÷Ïó£º

ÄÜÁ¦µÄ ¡°Ó¿ÏÖ¡±£ºÔÚ BootcampEVAL ÆÀ²â¼¯ÖУ¬£¬£¬£¬£¬Qwen2.5-32B Ä£×ӵį½¾ùÐÔÄÜʵÏÖÁË·­±¶Ê½ÔöÌí£¨´Ó 24.4 ÌáÉýÖÁ 59.5£©¡£¡£¡£¡£¡£¡£¡£¸üΪҪº¦µÄÊÇ£¬£¬£¬£¬£¬²¿·ÖÔÚµ¥Ê¹ÃüѵÁ·ÏÂÎÞ·¨½â¾öµÄÂß¼­Ê¹Ãü£¬£¬£¬£¬£¬ÔÚ¾­ÓÉ 500 ÓàÏî»ìÏýʹÃüѵÁ·ºó±äµÃ¿É½â¡£¡£¡£¡£¡£¡£¡£Õâ֤ʵÁËʹÃü¼äµÄÒþÐÔ¹ØÁªÄܹ»ÓÐÓÃÔöǿģ×ÓµÄ×ÛºÏÃ÷È·ÄÜÁ¦¡£¡£¡£¡£¡£¡£¡£Ê¹ÃüÀ©Õ¹¶¨ÂÉ£ºÊµÑéÊý¾ÝÏÔʾ£¬£¬£¬£¬£¬µ±Ê¹ÃüÀàÐÍÊýÄ¿´Ó 8 ÖÖÀ©Õ¹ÖÁ 512 ÖÖʱ£¬£¬£¬£¬£¬Ä£×ÓÐÔÄÜ·ºÆðÒ»Á¬ÉÏÉýÇ÷ÊÆ¡£¡£¡£¡£¡£¡£¡£ÕâһЧ¹û֤ʵÁËÓëʹÃüÊýÄ¿ÔöÌíÏà¹ØµÄ¹æÄ £» £»£»£»£»£»¯¶¨ÂÉÕæÊµ±£´æ£¬£¬£¬£¬£¬ÎªÎ´À´´ó¹æÄ£ÑµÁ·ÌṩÁËÀíÂÛÒÀ¾Ý¡£¡£¡£¡£¡£¡£¡£

ÂäµØÎ¬¶È£ºSimpleVLA-RL Óë¾ßÉíÖÇÄÜÑݽø

½ø»¯µÄÖÕ¾Ö£¬£¬£¬£¬£¬ÊǻعéÎïÀíÌìÏ¡£¡£¡£¡£¡£¡£¡£Ä¿½ñ¾ßÉíÖÇÄÜÃæÁٵĽ¹µãÆ¿¾±ÊÇÊý¾ÝØÑ·¦£º»úеÈËÑÝʾÊý¾Ý»ñÈ¡±¾Ç®¼«¸ß£¬£¬£¬£¬£¬ÇÒ´¿´âÀ©´ó¼àÊÓ΢µ÷£¨SFT£©¹æÄ£ÃæÁٱ߼ÊÐ§ÒæµÝ¼õ¡£¡£¡£¡£¡£¡£¡£ÎÒÃÇÒÔΪ£¬£¬£¬£¬£¬Ç¿»¯Ñ§Ï°£¨RL£©ÒÀ¸½ÆäÍ»ÆÆÑÝʾÊý¾Ý¾ÖÏÞµÄ̽Ë÷ÄÜÁ¦£¬£¬£¬£¬£¬ÍŽá¼òÆÓµÄ¶þÔª½±Àø£¨ÀÖ³É / ʧ°Ü£©£¬£¬£¬£¬£¬×ãÒÔ³ÉΪ½â¾öÕâÒ»ÎÊÌâµÄÔ¿³×¡£¡£¡£¡£¡£¡£¡£

»ùÓÚ´Ë£¬£¬£¬£¬£¬ÎÒÃÇÌá³öÁ˼«¶ËÊý¾ÝϡȱÇéÐÎϵÄÔÚÏßÇ¿»¯Ñ§Ï°¿ò¼Ü ¡ª¡ªSimpleVLA-RL?¡£¡£¡£¡£¡£¡£¡£¸Ã¿ò¼Ü»ùÓÚÊÓ¾õ - ÓïÑÔ - Ðж¯£¨VLA£©Ä£×Ó£¬£¬£¬£¬£¬ÍŽá GRPO ÓÅ»¯Ä¿µÄ£¬£¬£¬£¬£¬²¢Í¨¹ý²¢ÐжàÇéÐÎäÖȾÊÖÒÕÖ§³Ö½»»¥Ê½¹ì¼£²ÉÑù¡£¡£¡£¡£¡£¡£¡£

¼«¶ËÊý¾ÝϡȱÇéÐÎϵÄÔÚÏßÇ¿»¯Ñ§Ï°¿ò¼Ü SimpleVLA-RL

ʵÑéЧ¹ûÇ㸲Á˶ÔÊý¾ÝЧÂʵĹŰåÈÏÖª£º

³¬¸ßÊý¾ÝЧÂÊ£º½öÐè ¡°µ¥¹ì¼£¡± ¼àÊÓ΢µ÷ÍŽá RL£¬£¬£¬£¬£¬¼´¿ÉʵÏÖ 96.9% µÄÀÖ³ÉÂÊ£¬£¬£¬£¬£¬ÐÔÄÜ·´¶øÓâÔ½ÁËÈ«¹ì¼£¼àÊÓ΢µ÷ £» £»£»£»£»£»Õ½ÂÔÓ¿ÏÖ£º»úеÈËͨ¹ý RL ×ÔÖ÷̽Ë÷³öÁË´Óδ±»ÑÝʾ¹ýµÄÈ«ÐÂÍÆ¿ØÕ½ÂÔ£¬£¬£¬£¬£¬Õ¹ÏÖ³öǿʢµÄ˳ӦÐÔ £» £»£»£»£»£»Sim-to-Real Í»ÆÆ£ºÔÚµþÍëµÈµä·¶²Ù×÷ʹÃüÖУ¬£¬£¬£¬£¬·ÂÕæµ½ÏÖʵµÄǨáãÀÖ³ÉÂÊÌáÉýÁË 21% £» £»£»£»£»£»³¤Ê±³ÌʹÃüÄÜÁ¦£ºÔÚ½üÆÚÂ䵨ÖУ¬£¬£¬£¬£¬¸Ã¼Æ»®ÔÚ³¤Ê±³ÌÁéÇɲÙ×÷ʹÃüÉÏ£¬£¬£¬£¬£¬ÊµÏÖÁËÏà¶ÔÐÔÄÜÌáÉý 300%£¬£¬£¬£¬£¬²¢Õ¹ÏÖ³öÁîÈ˾ªÏ²µÄ×ÔÖ÷»Ö¸´ÄÜÁ¦¡£¡£¡£¡£¡£¡£¡£

µÃÒæÓÚ SimpleVLA-RL£¬£¬£¬£¬£¬ÎÒÃǽöÓÃÉÙÉÙµÄÊý¾ÝÓëÅÌËã×ÊÔ´£¬£¬£¬£¬£¬±ãÈ¡µÃÁË¿ÉÓë Physical Intelligence ÍÅ¶Ó ¦Ð*0.6 Ä£×ӱȼçµÄÐÔÄÜÌåÏÖ¡£¡£¡£¡£¡£¡£¡£ÕâһЧ¹û±ê¼Ç×Å SAGE ¼Ü¹¹³¹µ×ÂòͨÁËÈÏÕæÍÆÀí¾öÒéµÄ ¡°´óÄÔ¡± ÓëÈÏÕæÖ´ÐÐÐж¯µÄ ¡°ÇûÌ塱£¬£¬£¬£¬£¬ÕæÕýʵÏÖÁËÖÇÄÜÌåÔÚÎïÀíÌìÏÂÖÐµÄ ¡°¾ßÉí»¯¡± Ñݽø¡£¡£¡£¡£¡£¡£¡£

¾­ÓɽüÁ½ÄêµÄÔúʵ̽Ë÷£¬£¬£¬£¬£¬SAGE ¼Ü¹¹ÒÑ¿çÔ½ÀíÂÛ¹¹Ïë½×¶Î£¬£¬£¬£¬£¬Íê³ÉÁËȫջÑéÖ¤¡£¡£¡£¡£¡£¡£¡£ÔÚ»ù´¡²ã£¬£¬£¬£¬£¬MemoryDecoder ʵÏÖÁËÓ°ÏóÓëÅÌËãµÄ½á¹¹ÐÔ½âñî £» £»£»£»£»£»ÔÚÈںϲ㣬£¬£¬£¬£¬PRIME Óë FlowRL ¹¥¿ËÁ˼àÊÓϡȱÓëÍÆÀí¼òµ¥ÐÔµÄÄÑÌâ £» £»£»£»£»£»ÔÚ½ø»¯²ã£¬£¬£¬£¬£¬TTRL¡¢InternBootcamp Óë SimpleVLA-RL ¹¹½¨ÁË´Ó²âÊÔʱǿ»¯µ½ ¡°¾ßÉí»¯¡± ÑݽøµÄ±Õ»·¡£¡£¡£¡£¡£¡£¡£

·¶Ê½¸ïÃü£º´Ó AI4S µ½ AGI4S

Ö»¹ÜÒÔ AlphaFold Ϊ´ú±íµÄ AI for Science£¨AI4S£©ÊÖÒÕÔÚÂѰ×ÖÊÕÛµþ¡¢ÆøÏóÕ¹ÍûµÈÌØ¶¨ÁìÓòÈ¡µÃÁËÀï³Ì±®Ê½³É¼¨£¬£¬£¬£¬£¬µ«½üÆÚ¡¶Nature¡·½ÒÏþµÄÑо¿Ö¸³ö£¬£¬£¬£¬£¬Ì«¹ýÒÀÀµÏÖÓÐÉî¶Èѧϰģ×Ó¿ÉÄܾÖÏÞÐÂ֪ʶµÄ̽Ë÷½çÏߣ¬£¬£¬£¬£¬ÉõÖÁÔÚijÖÖˮƽÉÏ×è°­Á¢Òì¡£¡£¡£¡£¡£¡£¡£ÕâÓ¡Ö¤ÁËPTÊÓѶ(ÖйúÇø)¹ÙÍø½¹µã¿´·¨£ºÉÃÓŵãÀíÊý¾Ý¸»×ã¡¢½ç˵Ã÷ȷʹÃüµÄ¹Å°åÉî¶Èѧϰ£¬£¬£¬£¬£¬Èô½ö×÷Ϊ¹¤¾ß±£´æ£¬£¬£¬£¬£¬ÄÑÒÔÓ¦¶Ô¿ÆÑ§·¢Ã÷ÖÐ ¡°Î´ÖªµÄδ֪¡±¡£¡£¡£¡£¡£¡£¡£

ϵͳÐÔµÄÆÀ¹À½øÒ»²½Õ¹ÏÖÁËÄ¿½ñÇ°ÑØÄ£×ӵĶ̰塣¡£¡£¡£¡£¡£¡£ÎÒÃÇÍŽáÀ´×Ô 10 ¸ö²î±ð¿ÆÑ§ÁìÓòµÄ 100 λ¿ÆÑ§¼ÒÉè¼ÆÁËÆÀ¹Àϵͳ£¬£¬£¬£¬£¬Ð§¹ûÏÔʾ£ºÇ°ÑØÄ£×ÓÔÚͨÓÿÆÑ§ÍÆÀíʹÃüÖе÷ֿɴï 50 ·Ö£¨Âú·Ö 100£©£¬£¬£¬£¬£¬µ«ÔÚÖÖÖÖ×¨ÒµÍÆÀíʹÃü£¨ÈçרÏîÎÄÏ×¼ìË÷¡¢ÏêϸʵÑ鼯»®Éè¼Æ£©ÖУ¬£¬£¬£¬£¬µÃ·ÖÖè½µÖÁ 15-30 ·Ö¡£¡£¡£¡£¡£¡£¡£

ÕâÖÖÏÔ×ÅµÄ ¡°Ä¾Í°Ð§Ó¦¡± Åú×¢£¬£¬£¬£¬£¬¿ÆÑ§·¢Ã÷È«ÖÜÆÚµÄЧÄÜÕýÊÜÖÆÓÚ×¨ÒµÍÆÀíÄÜÁ¦µÄ×Èõ»·½Ú¡£¡£¡£¡£¡£¡£¡£Òò´Ë£¬£¬£¬£¬£¬ÕûºÏͨÓÃÍÆÀíÓëרҵÄÜÁ¦£¬£¬£¬£¬£¬½ø¶øÍƶ¯¿ÆÑ§ÖÇÄÜ´Ó AI4S Ïò AGI4S µü´ú³ÉΪһ¶¨Ñ¡Ôñ¡£¡£¡£¡£¡£¡£¡£

Ñо¿Åú×¢£¬£¬£¬£¬£¬Ä¿½ñËùÓÐÇ°ÑØÄ£×ӵĿÆÑ§ÄÜÁ¦¾ùÏÔȱ·¦

´Ó AI4S ÂõÏò AGI4S£¬£¬£¬£¬£¬ÕâÒ»Éý¼¶Ö¼ÔÚÍÆ¶¯Ñо¿Õß¡¢Ñо¿¹¤¾ßÓëÑо¿¹¤¾ßµÄЭͬÑݽø¡£¡£¡£¡£¡£¡£¡£Í¨¹ý AGI Ôö½øÈýÕßÏ໥×÷Óá¢Ð­Í¬Ñݽø¡¢ÂÝÐýʽÉÏÉý£¬£¬£¬£¬£¬½«´´Á¢³öÕæÕý¡°¸ïÃüµÄ¹¤¾ß¡±£¬£¬£¬£¬£¬Íƶ¯¿ÆÑз¶Ê½Àå¸ï?¡£¡£¡£¡£¡£¡£¡£

´Ó AI4S 1.0 µ½ AI4S 2.0£¨AGI4S£©

Intern-S1£ºÃæÏò¿ÆÑ§µÄ¿ÉÉî¶Èרҵ»¯Í¨ÓÃÄ£×Ó

ÎªÍ»ÆÆÉÏÊöÆ¿¾±£¬£¬£¬£¬£¬ÎÒÃÇÑз¢ÁË ¡°ÊéÉú¡± ¿ÆÑ§¶àģ̬´óÄ£×Ó£¨Intern-S1£©?¡£¡£¡£¡£¡£¡£¡£×÷Ϊ SAGE ¼Ü¹¹ÔÚ¿ÆÑ§ÁìÓòµÄ¼¯ÖÐÌåÏÖ£¬£¬£¬£¬£¬Intern-S1 Ö¼ÔÚ¹¹½¨Ò»¸ö¼È¾ß±¸Ç¿Ê¢Í¨ÓÃÄÜÁ¦£¬£¬£¬£¬£¬ÓÖÄÜÃ÷È·ÖØ´ó¿ÆÑ§Êý¾ÝµÄ ¡°¿ÉÉî¶Èרҵ»¯Í¨²Å¡±¡£¡£¡£¡£¡£¡£¡£ÆäÔÚÈý¸ö²ãÃæ¾ÙÐÐÁËÉî¶ÈÁ¢Ò죺

»ù´¡²ã£¨Êý¾ÝÊÊÅ䣩£ºÕë¶Ô¿ÆÑ§Êý¾ÝµÄ¶àģ̬Òì¹¹ÐÔ£¬£¬£¬£¬£¬Ìá³öÁË¿ÆÑ§×¨Óüܹ¹¡£¡£¡£¡£¡£¡£¡£½ÓÄɶ¯Ì¬·Ö´ÊÆ÷ÓëרÓñàÂëÆ÷£¬£¬£¬£¬£¬Ô­ÉúÖ§³Ö DNA ÐòÁС¢ÂѰ×Öʽṹ¡¢Ê±¼äÐòÁÐµÈ 10 ÓàÖÖģ̬¡£¡£¡£¡£¡£¡£¡£Ïà½ÏÓÚ GPT-OSS µÈͨÓÃÄ£×Ó£¬£¬£¬£¬£¬ÆäÔÚ¿ÆÑ§Êý¾ÝÉϵÄѹËõÂÊÌáÉýÁË 1.7 ±¶£¬£¬£¬£¬£¬²¢»ùÓÚ 2.5 ÍòÒÚ¸ßÖÊÁ¿¿ÆÑ§ Token ¾ÙÐÐÁËԤѵÁ·¡£¡£¡£¡£¡£¡£¡£Èںϲ㣨»ìÏý½±Àø£©£º¹¹½¨ÁË»ìÏý½±Àø¿ò¼Ü£¨MoR£©£¬£¬£¬£¬£¬½«¶àÖÖÇ¿»¯Ñ§Ï°Ëã·¨ÓëìØ»úÖÆÕûºÏ¡£¡£¡£¡£¡£¡£¡£¸Ã¿ò¼ÜƽºâÁËÅÌËã¡¢ÍÆÀí¡¢ÊµÑéÉè¼ÆµÈ²î±ðÊÖÒÕËùÐèµÄ½±ÀøÐźÅ£¬£¬£¬£¬£¬ÓÐÓûº½âÁËÌØ¶¨Ê¹Ãü¹ýÄâºÏÎÊÌ⣬£¬£¬£¬£¬ÔöÇ¿ÁËÄ£×ÓÔÚ¿çÁìÓòÖØ´óÍÆÀíÖеķº»¯ÄÜÁ¦¡£¡£¡£¡£¡£¡£¡£½ø»¯²ã£¨½»»¥×¨¾«£©£ºÒÀÍÐ InternBootCamp ¿ò¼Ü£¬£¬£¬£¬£¬Ä£×ÓÔÚ³¬ 1000 ÏîרҵʹÃü£¨ÈçÄæºÏÉíÆÊÎö£©ÖÐÓëÄ£ÄâÆ÷¾ÙÐн»»¥Ñ§Ï°£¬£¬£¬£¬£¬ÊµÏÖÁË´ó¹æÄ£µÄʹÃüר¾«¡£¡£¡£¡£¡£¡£¡£

²âÆÀЧ¹ûÏÔʾ£¬£¬£¬£¬£¬Intern-S1 ÔÚͨÓÃÄÜÁ¦ÉÏ¶ÔÆë SOTA ¿ªÔ´Ä£×Ó£¬£¬£¬£¬£¬¶øÔÚº­¸Ç»¯Ñ§¡¢ÉúÎï¡¢ÖÊÁÏµÈ 9 ´óÁìÓòµÄ¿ÆÑ§ÐÔÄÜÉÏ£¬£¬£¬£¬£¬ÖÜÈ«ÓâÔ½Á˰üÀ¨ GPT-5 ºÍ Grok-4 ÔÚÄڵĶ¥¼â±ÕÔ´Ä£×Ó¡£¡£¡£¡£¡£¡£¡£

Intern-Discovery£ºÈ«Á÷³Ì¿ÆÑ§ÖÇÄÜϵһÇÐ

ÈôÊÇ˵ Intern-S1 ÊÇ¿ÆÑ§´óÄÔ£¬£¬£¬£¬£¬ÄÇô Intern-Discovery ÔòÊǾ߱¸Ðж¯Á¦µÄ¿ÆÑ§ÖÇÄÜÌå¡£¡£¡£¡£¡£¡£¡£¸Ãƽ̨¹¹½¨ÁËÒ»¸ö½« Intern-S1 Ó뺣Á¿Êý¾Ý¡¢2000 + רҵ¹¤¾ß¼°ÊªÊµÑéÊÒÑéÖ¤ÇéÐÎÉî¶ÈÈںϵÄÖÇÄÜϵһÇУ¬£¬£¬£¬£¬ÊµÏÖÁË´Ó¼ÙÉèÌìÉúµ½ÊµÑéÑéÖ¤µÄ±Õ»·¡£¡£¡£¡£¡£¡£¡£

Intern-Discovery µÄ½¹µãÂß¼­ÔÚÓÚ½¨Éè ¡°ÖÇÄÜÌåÌìÉú¡± Óë ¡°ÖÇÄÜÌåÑéÖ¤¡± µÄË«ÏòÑ­»·£ºÇ°Õß×Ô¶¯¶´²ìÕ÷Ïó¡¢Ìá³ö¼ÙÉè²¢Éè¼ÆÊµÑé £» £»£»£»£»£»ºóÕßͨ¹ý·ÂÕæÓëÎïÀíʵÑéÑéÖ¤¼ÙÉ裬£¬£¬£¬£¬²¢½«·´Ïì»Ø´«ÒÔÐÞÕýÈÏÖª¡£¡£¡£¡£¡£¡£¡£

Ϊ֧³ÖÕâÒ»ÖØ´óÁ÷³Ì£¬£¬£¬£¬£¬ÏµÍ³ÒýÈëÁËÁ½´ó¸Åº¦Ö§Öù£º

¿ÆÑ§ÖÇÄÜÉÏÏÂÎÄЭÒ飨SCP£©?£ºÕë¶ÔÏÖÓÐ MCP ЭÒéÔÚ¿ÆÑ§×ÊÔ´ÕûºÏÉϵÄȱ·¦£¬£¬£¬£¬£¬SCP ½ç˵ÁËÁìÓòÌØ¶¨µÄ½á¹¹ÓëЭµ÷»úÖÆ£¬£¬£¬£¬£¬ÊµÏÖÁ˶ÔÊý¾Ý¼¯¡¢ÊªÊµÑéÊÒ×°±¸¼°ÖØ´óÊÂÇéÁ÷µÄ±ê×¼»¯µ÷ÀíÓëÈ«ÉúÃüÖÜÆÚÖÎÀí¡£¡£¡£¡£¡£¡£¡£·Ö²ãÓ°ÏóÄ£¿£¿£¿£¿£¿£¿£¿é£ºÍ¨¹ýÕ½ÂÔ³ÌÐòÓ°Ïó£¨SPM£©¡¢Ê¹ÃüÇé¾°Ó°Ïó£¨TEM£©ÓëÓïÒå֪ʶӰÏó£¨SKM£©µÄЭͬ£¬£¬£¬£¬£¬ÏµÍ³Äܹ»³Áµí¸ß½×Ñо¿Ä£Ê½¡¢¼Í¼ʵÑéϸ½Ú²¢ÕûºÏºã¾Ã֪ʶ£¬£¬£¬£¬£¬´Ó¶øÔÚÒ»Á¬µü´úÖÐ×èÖ¹Âß¼­»Ã¾õ¡£¡£¡£¡£¡£¡£¡£

°¸Àýʵ֤£ºÖØËÜ¿ÆÑ§·¢Ã÷Á÷³Ì

Intern-Discovery ÒÑÔÚÌìÆø¿ÆÑ§ÓëÉúÎïҽѧÁìÓòÕ¹ÏÖ³ö ¡°¸ïÃüÐÔ¹¤¾ß¡± µÄDZÁ¦¡£¡£¡£¡£¡£¡£¡£

ÔÚÌìÆø¿ÆÑ§ÁìÓò£¬£¬£¬£¬£¬ÃæÁÙ½µË®Õ¹ÍûÖм«¶ËÖØ´óµÄ·ÇÏßÐÔ½»»¥£¬£¬£¬£¬£¬Intern-Discovery ×ÔÖ÷ŲÓà 30 ÓàÖÖ¹¤¾ß£¬£¬£¬£¬£¬ÆÊÎöÁË 20 ÄêµÄ¶àģ̬Êý¾Ý¡£¡£¡£¡£¡£¡£¡£ËüдÁË 4000 ¶àÐÐרҵ´úÂ룬£¬£¬£¬£¬Àֳɷ¢Ã÷Á˱»ÈËÀàר¼ÒºöÂÔµÄË®ÆûÓ붯Á¦Ïî¹ØÁª£¬£¬£¬£¬£¬²¢ÍƵ¼³öÒ»¸ö¾«Á·µÄÐÂÐÍÏÔʽ·ÇÏßÐÔ·½³Ì¡£¡£¡£¡£¡£¡£¡£¸Ã·½³Ì²»µ«ÐÎʽÓÅÑž«Á·£¬£¬£¬£¬£¬ÇÒÏÔÖøÌáÉýÁËÄ£Ä⾫¶È£¬£¬£¬£¬£¬ÓÐÓÃÐÞÕýÁ˺ã¾Ã±£´æµÄϵͳÐÔÎó²î£¬£¬£¬£¬£¬Ö¤ÊµÎúÖÇÄÜÌåÔÚÀíÂÛ¹¹½¨²ãÃæµÄ´´Á¢Á¦?¡£¡£¡£¡£¡£¡£¡£

Intern-Discovery ÔÚÌìÆø¿ÆÑ§µÄÓ¦Óð¸Àý

ÔÚÉúÎïҽѧÁìÓò£¬£¬£¬£¬£¬ÐéÄâ¼²²¡ÉúÎïѧ¼Ò ¡°ÔªÉú¡± ͨ¹ýÄ£ÄâÈËÀà¿ÆÑ§¼ÒµÄÍ·ÄÔÄ£°å£¬£¬£¬£¬£¬ÕûºÏÒÅ´«Ñ§¡¢ÂѰ×ÖÊ×éѧ¼°ÁÙ´²ÎÄÏ׵ȶàÔ´Êý¾Ý¡£¡£¡£¡£¡£¡£¡£¼´±ãÔÚÊý¾ÝÏ£º±Ìõ¼þÏ£¬£¬£¬£¬£¬ËüÈÔÀֳɷ¢Ã÷²¢ÑéÖ¤Á˾ßÓиßÁÙ´²Ç±Á¦µÄÒþ²Ø°Ðµã£¬£¬£¬£¬£¬Õ¹Ê¾ÁË´ÓÊý¾Ýµ½»úÖÆ¡¢´Ó¼Ù˵µ½ÑéÖ¤µÄÈ«Á÷³ÌÖÇÄÜ»¯ÄÜÁ¦¡£¡£¡£¡£¡£¡£¡£

Intern-Discovery ÔÚÉúÎïҽѧµÄÓ¦Óð¸Àý

´Ó Intern-S1 µÄµ×²ãÍÆÀíÍ»ÆÆµ½ Intern-Discovery µÄϵͳ¼¶Ó¦Ó㬣¬£¬£¬£¬ÎÒÃÇÕýÖð²½¹¹½¨ÆðÒ»Ì×ÁýÕÖ¿ÆÑ§·¢Ã÷È«ÖÜÆÚµÄ AGI4S »ù´¡ÉèÊ©¡£¡£¡£¡£¡£¡£¡£Õâ²»µ«Êǹ¤¾ßµÄˢУ¬£¬£¬£¬£¬¸üÊÇ¿ÆÑз¶Ê½µÄÖØËÜ ¡ª¡ª ÈÃÈ˹¤ÖÇÄÜÕæÕý³ÉÎªÍÆ¶¯¿ÆÑ§½çÏßÍØÕ¹µÄÏàÖúͬ°é¡£¡£¡£¡£¡£¡£¡£

Ðж¯ÕÙ»½£º¹²ÍØÐÂÌìÏÂÀ¶Í¼

×ÛÉÏËùÊö£¬£¬£¬£¬£¬ÎÒÃÇÕý´¦ÔÚʵÏÖ AGI µÄǰϦ£¬£¬£¬£¬£¬ÈôAGI = ͨרÈںϣ¨Specialized Generalist£©£¬£¬£¬£¬£¬Ôò¿ÉÉî¶Èרҵ»¯µÄͨÓÃÄ£×Ó£¨Specializable Generalist£©ÊÇʵÏÖ AGI µÄ¿ÉÐз¾¶£¬£¬£¬£¬£¬¶ø¡°ÖÇÕß¡±SAGE µÄÈý²ãÊÖÒÕ¿ò¼ÜÕýÊÇÇý¶¯ºóÕßÉú³¤µÄ½¹µã¼Ü¹¹¡£¡£¡£¡£¡£¡£¡£

ÏÂÒ»¸öÇ°ÑØÕóµØÊÇ¿ÆÑ§·¢Ã÷ ¡ª¡ª Ëü¼ÈÊÇÍÆÀíÖÇÄܵÄ×îÖÕÊÔÁ¶³¡£¬£¬£¬£¬£¬Ò²ÊÇ ¡°Í¨×¨Èںϡ± µÄÑéÖ¤Îę̀£¬£¬£¬£¬£¬´ó¹æÄ£ÍÆÀí½«¸³ÄÜ¿ÆÑ§·¢Ã÷£¬£¬£¬£¬£¬¿ÆÑ§·¢Ã÷Òཫ·´²¸ÍÆÀíÄÜÁ¦µÄ½ø»¯¡£¡£¡£¡£¡£¡£¡£

Intern-S1 Óë Intern-Discovery ÊÇÂõÏò¸ÃÆ«ÏòµÄÊײ½Êµ¼ù£¬£¬£¬£¬£¬µ«ÕâÒ»Çнö½öÊdzõʼµÄ³ûÐΡ£¡£¡£¡£¡£¡£¡£ÈôÊǽ«¡°ÖÇÕß¡±SAGE ¼Ü¹¹±È×÷Ò»ÕÅÐÂÌìϵĵØÍ¼£¬£¬£¬£¬£¬ÎÒÃÇÏÖÔÚÒѽ¨ÉèÁËºÜºÃµÄÆðÔ´ÑéÖ¤ÓëÐí¶à¼â±øÇ°ÉÚÕ¾£¬£¬£¬£¬£¬µ«ÕâÕŵØÍ¼ÉÏÈÔ±£´æÁÉÀ«µÄ ¡°¿ÕÈ±ÇøÓò¡±¡£¡£¡£¡£¡£¡£¡£

¼Ü¹¹ÒѾ­Í£µ±£¬£¬£¬£¬£¬µ«»­¾íÈÔ±£´æ´óƬÁô°×¡£¡£¡£¡£¡£¡£¡£ÈôÊÇÕâЩÆðÔ´Ï£Íû¼¤ÆðÁËÄãµÄÐËȤ£¬£¬£¬£¬£¬ÎÒÔ¼ÇëÄãÉîÈëÔĶÁPTÊÓѶ(ÖйúÇø)¹ÙÍøÂÛÎÄÓë´úÂë ¡ª¡ª ËüÃǶ¼ÊÇ¿ªÔ´µÄ¡£¡£¡£¡£¡£¡£¡£µ«¸üÖ÷ÒªµÄÊÇ£¬£¬£¬£¬£¬ÎÒÔ¼Çë־ͬ־ºÏÕßÓëÎÒÃÇһͬÌî²¹ÕâЩ¿Õȱ£¬£¬£¬£¬£¬ÅäºÏ¹¹½¨ÍêÕûµÄÀ¶Í¼¡£¡£¡£¡£¡£¡£¡£

лл£¡

±¾´Î±¨¸æ½¹µãÒªµã×ܽá

²Î¿¼ÎÄÏ×

¢Ù Shanghai Artificial Intelligence Laboratory. Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows [J]. arXiv preprint arXiv:2512.16969v1, 2025.

¢Ú Vaswani A, et al. Attention is all you need [C]// Advances in neural information processing systems, 2017, 30.

¢Û Zhang K, Qi B, Zhou B. Towards building specialized generalist ai with system 1 and system 2 fusion [J]. arXiv preprint arXiv:2407.08642, 2024.

¢Ü Qi B, Zhang K, Tian K, ..., Zhou B. Large language models as biomedical hypothesis generators: a comprehensive evaluation [C]. COLM, 2024.

¢Ý Zhou B. Building AGI through Specialized Generalist AI: pathways and key issues [J]. Communications of CCF, 2025, 21 (1): 54-62.

¢Þ Cao J, Wang J, Wei R, ..., Zhou B, Lin Z. Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Models [J]. arXiv preprint arXiv:2508.09874, 2025.

¢ß Zhang K, Zuo Y, He B, ..., Zhou B. A survey of reinforcement learning for large reasoning models [J]. arXiv preprint arXiv:2509.08827, 2025.

¢à Cui G, Yuan L, Wang Z, ..., Zhou B, Ding N. Process Reinforcement through Implicit Rewards [J]. arXiv preprint arXiv:2502.01456, 2025.

¢á Cui G, Zhang Y, Chen J, ..., Zhou B, Ding N. The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models [J]. arXiv preprint arXiv:2505.22617, 2025.

¢â Zhu X, Cheng D, Zhang D, ..., Zhou B, Mei H, Lin Z. FlowRL: Matching reward distributions for LLM reasoning [J]. arXiv preprint arXiv:2509.15207, 2025.

? Zuo Y, Zhang K, Sheng L, ..., Ding N, Zhou B. TTRL: Test-Time Reinforcement Learning [C]// NeurIPS, 2025.

? Li P, Ye J, Chen Y, ..., Zhou B, Chen K. InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling [J]. arXiv preprint arXiv:2508.08636, 2025.

? Li H, Zuo Y, Yu J, ..., Zhou B, Ding N. SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning [J]. arXiv preprint arXiv:2509.09674, 2025.

? Zhou B, Ding N, Bai L, Zhou H. Advancing AI for science: From the revolution of tools to the tools for revolution [J]. AI Open, 2025, 6: 323-328.

? Shanghai AI Laboratory. INTERN-S1: A SCIENTIFICMULTIMODAL FOUNDATION MODEL [J]. arXiv preprint arXiv:2508.15763, 2025.

? Jiang Y, Lou W, Wang L, ..., Zhou B. SCP: Accelerating Discovery with a Global Web of Autonomous Scientific Agents [J]. arXiv preprint arXiv:2512.24189, 2025.

? Guo Z, Wang ÔÆÄÏµá³ØÂÌɫʳÎïÓÐÏÞ¹«Ë¾J, Ling F, ..., Zhou B, Bai L. A Self-Evolving AI Agent System for Climate Science [J]. arXiv preprint arXiv:2507.17311v3, 2025.