Skip to content

ABot-Earth 0.5: Generative 3D Earth Model

https://huggingface.co/papers/2606.09967

https://arxiv.org/abs/2606.09967

https://abot-earth.amap.com/

https://github.com/amap-cvlab/ABot-Earth-0.5

Alibaba AMAP CV Lab

ABot-Earth 0.5:生成式 3D 地球模型

We present ABot-Earth 0.5, a generative 3D framework designed to synthesize vast, seamless 3D environments from ubiquitous, geospatially referenced satellite imagery. To achieve this, we propose a novel generative model formulated directly with the 3D Gaussian Splatting (3DGS) representation. The model is trained on a diverse corpus of existing real-world urban reconstructions, learning to generate realistic geometry and textures. At inference, it synthesizes novel 3D scenes conditioned solely on satellite imagery at a scalable rate of under 10 minutes per square kilometer, while demonstrating exceptional realism. The framework is designed for accessibility, with integrated hierarchical level-of-detail (LOD) structures that permit real-time, interactive visualization on web-based map engines. This high-fidelity simulation sandbox effectively mitigates the sim-to-real domain gap, enabling critical downstream Embodied AI applications like closed-loop UAV navigation. By providing an ultra-low-cost and high-efficiency solution, ABot-Earth 0.5 significantly lowers the technical and financial barriers to large-scale 3D reconstruction and empowers the future of global digital earth visualization.

我们提出 ABot-Earth 0.5,一个从随处可得且带有地理空间参照的卫星影像中合成广阔、无缝 3D 环境的生成式 3D 框架。为实现这一目标,我们提出一种直接基于 3D 高斯泼溅(3DGS)表示构建的新型生成模型。该模型在多样化的现有真实城市重建语料上训练,从中学习生成逼真的几何结构与纹理。在推理时,它仅以卫星影像为条件合成新的 3D 场景,每平方公里耗时不到 10 分钟,具有良好的规模扩展能力,同时呈现出出色的真实感。该框架注重易用性,集成了层级细节层次(LOD)结构,支持在网页地图引擎上进行实时交互式可视化。这一高保真仿真沙盒有效缩小了仿真到真实的领域差距,支持闭环无人机导航等关键下游具身智能应用。ABot-Earth 0.5 通过提供成本极低且效率很高的解决方案,显著降低了大规模 3D 重建的技术与资金门槛,并推动全球数字地球可视化的未来发展。


Looped World Models

https://huggingface.co/papers/2606.18208

https://arxiv.org/abs/2606.18208

FaceMind

循环式世界模型

Current world models face a fundamental tension: faithful long-horizon simulation demands deep computation, but deeper models are expensive to deploy and prone to compounding errors. We resolve this by introducing Looped World Models (LoopWM), which are the first looped architectures for world modelling. Our method iteratively refines latent environment states through a parameter-shared transformer block. This yield up to 100x parameter efficiency over conventional approaches with adaptive computation that automatically scales depth to match the complexity of each prediction step. Orthogonal to scaling model size and training data, LoopWM establishes iterative latent depth as a new scaling axis for world simulation, which might significantly push the community forward.

当前的世界模型面临一种根本矛盾:忠实的长程仿真需要深层计算,但更深的模型部署成本高昂,并且容易产生累积误差。我们通过提出循环式世界模型(LoopWM)解决这一问题,这是首种用于世界建模的循环架构。我们的方法通过参数共享的 Transformer 块反复细化潜在环境状态。它采用自适应计算,自动调整深度以匹配每个预测步骤的复杂度,相比传统方法可实现最高 100 倍的参数效率。LoopWM 与扩大模型规模和训练数据规模相互独立,将迭代式潜在深度确立为世界仿真的新扩展轴线,有望显著推动该领域发展。


Agents' Last Exam

https://huggingface.co/papers/2606.05405

https://arxiv.org/abs/2606.05405

https://agents-last-exam.org/

https://github.com/rdi-berkeley/agents-last-exam

UC Berkeley

智能体终极考试

Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. We argue that this gap is largely an evaluation problem: widely used benchmarks lack sustained performance measurement on real and economically valuable workflows. This paper introduces Agents' Last Exam (ALE), a benchmark designed to evaluate AI agents on long-horizon, economically valuable, real-world tasks with verifiable outcomes. Developed in collaboration with 250+ industry experts, ALE covers non-physical industries defined with reference to O*NET / SOC 2018 (the U.S. federal occupational taxonomy). It is organized around a task taxonomy with 55 subfields grouped into 13 industry clusters covering 1K+ tasks. Current results show that the hardest tier remains far from saturated: across mainstream harness and backbone configurations, the average full pass rate is 2.6%. ALE is designed as a living benchmark: its task pool grows continuously as new workflows and industries are onboarded. More broadly, ALE is intended not merely as another leaderboard, but as an instrument for closing the gap between benchmark success and GDP-relevant impact.

近期 AI 系统在许多基准上取得了优异成绩,但这些进展尚未转化为众多专业领域中具有经济意义的部署。我们认为,这一差距很大程度上是评测问题:广泛使用的基准缺少对真实且具有经济价值的工作流进行持续性能衡量。本文提出智能体终极考试(ALE),一个用于评测 AI 智能体完成长程、具有经济价值且结果可验证的真实世界任务的基准。ALE 与 250 多名行业专家合作开发,覆盖参照美国联邦职业分类体系 O*NET / SOC 2018 定义的非实体行业。其任务分类体系包含 55 个子领域,归入 13 个行业集群,共覆盖 1,000 多项任务。当前结果表明,最难层级距离饱和仍十分遥远:在主流运行框架和骨干模型配置中,平均完全通过率仅为 2.6%。ALE 被设计为动态发展的基准,随着新工作流和行业接入,其任务池会持续增长。更广泛地说,ALE 的目标不只是成为另一个排行榜,而是作为一种工具,弥合基准成功与影响国内生产总值的实际成效之间的差距。


Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs

https://huggingface.co/papers/2605.30611

https://arxiv.org/abs/2605.30611

https://github.com/HaozheZhao/Crafter


Crafter:面向多样化输入可编辑科学图生成的多智能体运行框架

Scientific figures are among the most effective means of communicating complex research ideas, yet producing publication-quality illustrations remains one of the most labor-intensive parts of paper preparation. Existing automated systems each target a single figure type under text-only input, leaving the diversity of types and conditions researchers actually use unaddressed; their raster outputs further cannot be locally revised. Because scientific figures are structured compositions of discrete semantic components, the localized errors generators produce on such layouts demand not a stronger backbone but a harness. We instantiate this harness in two complementary systems: Crafter, a multi-agent harness for figure generation that generalizes across figure types and input conditions without architectural changes, and CraftEditor, which applies the same pattern to convert raster outputs into editable SVGs. Moreover, we introduce CraftBench, a benchmark spanning three figure types and four input conditions with human quality annotation. Experiments show that Crafter substantially outperforms both standalone generators and the agentic baseline on PaperBanana-Bench and CraftBench, with ablations confirming each component's independent contribution; CraftEditor faithfully converts outputs into editable SVGs that surpass all baselines. Our code and benchmark are available at https://github.com/HaozheZhao/Crafter.

科学图是传达复杂研究思想最有效的方式之一,但制作达到发表质量的插图仍是论文准备过程中最耗费人力的环节之一。现有自动化系统通常只在纯文本输入下针对单一图形类型,未能处理研究人员实际使用的多样图形类型与输入条件;其光栅输出也无法进行局部修改。科学图是由离散语义组件构成的结构化组合,因此生成器在这类布局上产生的局部错误所需要的并非更强的骨干模型,而是一个运行框架。我们通过两个互补系统实现这一框架:Crafter 是一个用于图形生成的多智能体运行框架,无需修改架构即可跨图形类型和输入条件泛化;CraftEditor 则采用相同模式,将光栅输出转换为可编辑的 SVG。我们还提出 CraftBench,一个涵盖三种图形类型和四种输入条件、带有人工质量标注的基准。实验表明,在 PaperBanana-Bench 和 CraftBench 上,Crafter 显著优于独立生成器和智能体基线,消融实验也验证了各组件的独立贡献;CraftEditor 能够忠实地将输出转换为可编辑 SVG,并超越所有基线。我们的代码和基准位于 https://github.com/HaozheZhao/Crafter。


On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters

https://huggingface.co/papers/2606.02437

https://arxiv.org/abs/2606.02437

Mind Lab

论 PEFT 的规模扩展:迈向万亿参数基础模型上的百万个个性化模型

Parameter-efficient fine-tuning (PEFT) is usually treated as a cheaper alternative to full fine-tuning. We study a broader role: small trainable adapters as persistent local state on top of strong shared foundation models. In this framing, the base model provides shared competence while adapters carry instance-specific behavior such as preferences, skills, tool habits, and memory-like updates. We organize the problem around three scaling axes: Scale Up, where stronger shared priors make small local updates more useful; Scale Down, where we study how small adapters can be while remaining reliable; and Scale Out, where many persistent adapted instances coexist. MinT provides one infrastructure example for managing adapter identity, revision, provenance, evaluation, and serving residency. Together, the results suggest that PEFT can be a compact substrate for persistent personal models rather than only a budget substitute for full fine-tuning.

参数高效微调(PEFT)通常被视为完整微调的一种低成本替代方案。我们研究它更广泛的作用:让小型可训练适配器充当强大共享基础模型之上的持久局部状态。在这一框架中,基础模型提供共享能力,而适配器承载偏好、技能、工具使用习惯和类似记忆的更新等实例特定行为。我们围绕三条扩展轴线组织这一问题:纵向扩展中,更强的共享先验使小型局部更新更加有效;向下缩减中,我们研究适配器在保持可靠的前提下可以缩小到何种程度;横向扩展中,大量持久的适配实例同时存在。MinT 提供了一个基础设施示例,用于管理适配器身份、修订版本、来源、评测和服务驻留状态。综合来看,结果表明 PEFT 可以成为持久个性化模型的紧凑载体,而不仅是完整微调的低预算替代品。


JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence

https://huggingface.co/papers/2606.14777

https://arxiv.org/abs/2606.14777

https://joyai-vl-video-future-academy-jd.github.io/JoyAI-VL-Interaction/

https://github.com/jd-opensource/JoyAI-VL-Interaction

JD.com Open Source

JoyAI-VL-Interaction:实时视觉语言交互智能

Many moments in the real world do not wait for a user to ask. A fire starts on a security monitor, an expression flickers across a video call, or a product a viewer wants flashes by in a livestream. Yet today's large models remain mostly turn-based by design: they answer only when addressed, and even video-call apps that appear interactive still operate as question-answer systems, reacting only when polled or prompted. We argue for a different paradigm: a model that is present in the world like a person. It continuously watches what is happening now, decides on its own whether to speak or stay silent, interacts in real time, and delegates to a background model when the problem is hard. To advance interaction models and their adoption across domains, we make two fully open-sourced contributions. First, we release JoyAI-VL-Interaction, an 8B-scale, vision-first VL-interaction model. The model makes the response decision internally, choosing each second to stay silent, respond, or delegate to a background model, and it excels at vision-triggered responsiveness and time awareness. We pair it with a transferable training recipe, from which capabilities we never trained for emerge, such as guiding a shopper through changing app screens or improvising a lecture from a slide deck. Second, we release a complete, deployable system built around that model. The system streams any ongoing video into the model, making it genuinely present in the world. All other components are pluggable, including ASR/TTS modules, memory, visualization UI, and a background brain that can connect to any API or agent. Across six real-world scenarios, human raters prefer JoyAI-VL-Interaction over the in-app video-call assistants of Doubao and Gemini by a wide margin. To our knowledge, this is the first open, vision-driven interaction model released together with its training recipe, data, and complete deployable system.

真实世界中的许多时刻不会等待用户提问。安防监控中突然起火,视频通话中一闪而过的表情,或直播中迅速出现又消失的观众心仪商品,都要求系统主动反应。然而,当前的大模型在设计上仍以轮次交互为主:它们只有在被问及时才回答,即使看似具有交互性的视频通话应用,本质上仍是问答系统,只会在被查询或提示时作出反应。我们主张采用一种不同的范式:让模型像人一样存在于世界中。它持续观察当下正在发生的事情,自主决定开口还是保持沉默,进行实时交互,并在问题困难时将任务委托给后台模型。为推进交互模型及其跨领域应用,我们作出两项完全开源的贡献。首先,我们发布 JoyAI-VL-Interaction,一个 80 亿参数规模、以视觉为先的视觉语言交互模型。该模型在内部作出响应决策,每秒选择保持沉默、作出响应或委托后台模型,并且擅长由视觉触发的及时响应和时间感知。我们为其配套一种可迁移的训练方案,由此涌现出从未专门训练过的能力,例如引导购物者操作不断变化的应用界面,或根据一套幻灯片即兴授课。其次,我们发布一个围绕该模型构建的完整可部署系统。该系统将任何正在播放的视频流输入模型,使其真正存在于现实环境中。其他所有组件均可插拔,包括 ASR/TTS 模块、记忆、可视化界面,以及能够连接任意 API 或智能体的后台大脑。在六种真实世界场景中,人工评测者以明显优势更偏好 JoyAI-VL-Interaction,而不是豆包和 Gemini 应用内的视频通话助手。据我们所知,这是首个将训练方案、数据和完整可部署系统一并发布的开放式视觉驱动交互模型。


LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling

https://huggingface.co/papers/2606.18023

https://arxiv.org/abs/2606.18023


LoopCoder-v2:只循环一次,实现高效测试时计算扩展

Looped Transformers scale latent computation by repeatedly applying shared blocks, but sequential looping increases latency and KV-cache memory with the loop count. Parallel loop Transformers (PLT) alleviate this cost through cross-loop position offsets (CLP) and shared-KV gated sliding-window attention, making loop count a practical design choice. We therefore study PLT loop-count selection through a gain--cost view: an extra loop may refine representations, but CLP also introduces a positional mismatch at each loop boundary. We instantiate this study by training LoopCoder-v2, a family of 7B PLT coders with different loop counts, from scratch on 18T tokens, followed by matched instruction tuning and evaluation. Empirically, the two-loop variant delivers broad gains over the non-looped baseline across code generation, code reasoning, agentic software engineering, and tool-use benchmarks, improving SWE-bench Verified from 43.0 to 64.4 points and Multi-SWE from 14.0 to 31.0 points. In contrast, variants with three or more loops regress, revealing a strongly non-monotonic loop-count effect. Our diagnostics show that loop 2 provides the main productive refinement, while later loops yield diminishing, oscillatory updates and reduced representational diversity. Because the CLP-induced mismatch remains roughly fixed as refinement gains shrink, the offset cost increasingly dominates. This gain--cost trade-off explains PLT's saturation at two loops and provides diagnostics for loop-count selection.

循环式 Transformer 通过反复应用共享模块扩展潜在计算,但顺序循环会使延迟和 KV cache 内存随循环次数增加。并行循环 Transformer(PLT)通过跨循环位置偏移(CLP)和共享 KV 的门控滑动窗口注意力缓解这一成本,使循环次数成为一种切实可用的设计选择。因此,我们从收益与成本的视角研究 PLT 的循环次数选择:额外一次循环可能细化表示,但 CLP 也会在每个循环边界引入位置失配。我们从零开始在 18 万亿个 token 上训练 LoopCoder-v2,以此开展研究。LoopCoder-v2 是一组循环次数不同的 70 亿参数 PLT 代码模型,随后接受配置匹配的指令微调与评测。实验上,两次循环的变体在代码生成、代码推理、智能体软件工程和工具使用基准上均比无循环基线获得广泛提升,将 SWE-bench Verified 从 43.0 分提高到 64.4 分,将 Multi-SWE 从 14.0 分提高到 31.0 分。相比之下,循环三次或更多的变体出现性能退化,揭示出强烈非单调的循环次数效应。我们的诊断表明,第二次循环提供了主要的有效细化,后续循环则产生收益递减、振荡式更新和表示多样性下降。随着细化收益缩小,由 CLP 引起的失配大致保持不变,因此偏移成本愈发占据主导。这种收益与成本之间的权衡解释了 PLT 为何在两次循环时达到饱和,并为循环次数选择提供了诊断依据。