论文
arxiv:2609.18703
复制 Markdown
RayOrch:用于基础模型数据准备的 lineage 控制多粒度数据流的编程与执行
发布于 9月16日
·
提交者
马晓晨
于 9月28日
·
北京大学
作者:
马晓晨
,
孟子墨
,
梁俊竹
,
蒋有合
,
程跃
,
梁浩
,
曾博涵
,
李登春
,
马璐
,
赵正阳
,
王振豪
,
何润明
,
强美怡
,
关江涛
,
袁彬航
,
张文涛
摘要
为基础模型准备高质量训练数据需要可扩展的流水线,将异构文档和视频转换为结构化记录。此类流水线将每个父项扩展为有序且依赖于输入的子项序列,其数量可能呈现长尾分布。GPU 应在保留父级关系、子项顺序、完成状态和结果路由的同时,跨父项对子项进行批处理。现有系统要么在粗粒度作业背后隐藏并行性,要么暴露扁平记录,迫使应用程序管理 lineage 和重新分组。我们提出了 RayOrch,一种在整个执行过程中保留父子关系的编程模型和分布式执行引擎。程序声明有序的可变基数扩展和匹配的收集操作。编译器验证每一对,而运行时记录子项成员资格、直接父级、不可变序数和终端状态。每个调用的 FIFO 就绪队列跨父项批处理已就绪的子项。收集操作根据声明的成员资格和序数而非批处理边界或完成顺序来重建结果。一旦所有必需的子项变为终端状态,父项即可推进。类型化的父级作用域失败会抑制失败父级的未分发兄弟节点,同时允许不相关的父级继续运行。在 NVIDIA H20 GPU 上,当将 MinerU 从 4 张 GPU 扩展到 64 张 GPU 时,RayOrch 实现了 15.14 倍的加速;当将视频流水线从 8 张 GPU 扩展到 64 张 GPU 时,实现了 7.82 倍的加速。与 Ray Data 相比,它在 MinerU 上将端到端时间减少了 13.1%,与 Daft 相比减少了 29.0%;在 Docling 上,与 Ray Data 相比减少了 16.0%。代码可在 https://github.com/OpenDCAI/RayOrch 获取。
查看 arXiv 页面
查看 PDF
项目页面
GitHub
16
添加到合集
社区
Sunnyhaze
论文作者
论文提交者
大约 16 小时前
我们介绍了 RayOrch,这是一个用于多模态基础模型数据准备的分布式数据流水线系统。与 Ray Data 和 Daft 类似,它针对大规模数据处理。RayOrch 专为具有多个阶段和粒度级别的流水线而设计,例如将文档按页处理、将视频按帧处理。它追踪每个部分属于哪个输入,为 GPU 执行跨输入批处理就绪的工作,并在任务无序完成时以正确的顺序收集结果。在 MinerU 工作负载上,与 Ray Data 相比,它将端到端时间减少了 13.1%,与 Daft 相比减少了 29.0%。我们欢迎反馈和问题!
查看翻译
🔥
1
+
回复
researchstudio-bot
大约 13 小时前
这是来自 ResearchStudio 团队的自动消息。
我们为这篇论文创建了一个交互式 ResearchStudio Reel。它包括视觉海报、视频和博客,所有文件均可下载且格式可编辑。
打开 ResearchStudio Reel →
从 Hugging Face 下载所有文件
如果您觉得该 Reel 有帮助,请给此评论点赞!
想探索或为更多论文创建 Reels?请访问 ResearchStudio 演示。
查看翻译
回复
编辑
预览
通过在文本输入中拖放、粘贴或点击此处来上传图片、音频和视频。
评论
· 注册或登录以发表评论
赞成票
43
+31
在您的代理中获取此论文:
hf papers read 2609.18703
没有最新的 CLI?
引用此论文的模型
0
无链接此论文的模型
在模型的 README.md 中引用 arxiv.org/abs/2609.18703 以从本页面链接它。
引用此论文的数据集
0
无链接此论文的数据集
在数据集的 README.md 中引用 arxiv.org/abs/2609.18703,以便从本页面链接到它。
引用此论文的 Spaces
0
没有链接此论文的 Space
在 Space 的 README.md 中引用 arxiv.org/abs/2609.18703,以便从本页面链接到它。
包含此论文的数据集集合
3
代码
数据集集合
82 项
•
更新于约 11 小时前
•
7
数据集与数据处理
数据集集合
10 项
•
更新于约 11 小时前
1
数据集集合
6 项
•
更新于约 3 小时前
Papers
arxiv:2609.18703
Copy markdown
RayOrch: Programming and Executing Lineage-Controlled Multi-Grain Dataflows for Foundation-Model Data Preparation
Published on Sep 16
·
Submitted by
Xiaochen Ma
on Sep 28
·
Peking University
Authors:
Xiaochen Ma
,
Zimo Meng
,
Junzhu Liang
,
Youhe Jiang
,
Yue Cheng
,
Hao Liang
,
Bohan Zeng
,
Dengchun Li
,
Lu Ma
,
Zhengyang Zhao
,
Zhen Hao Wong
,
Runming He
,
Meiyi Qiang
,
Jiangtao Guan
,
Binhang Yuan
,
Wentao Zhang
Abstract
Preparing high quality training data for foundation models requires scalable pipelines that transform heterogeneous documents and videos into structured records. Such pipelines expand each parent item into an ordered and input dependent sequence of children, whose counts may be long tailed. GPUs should batch children across parents while preserving parent relationships, child order, completion status, and result routing. Existing systems either hide parallelism behind coarse grained jobs or expose flat records that force applications to manage lineage and regrouping. We present RayOrch, a programming model and distributed execution engine that preserves parent child relations throughout execution. Programs declare ordered variable cardinality expansions and matching gathers. The compiler validates each pair, while the runtime records child membership, immediate parents, immutable ordinals, and terminal states. Per Call FIFO Ready Queues batch ready children across parents. Gathers reconstruct results from declared membership and ordinals rather than batch boundaries or completion order. Parents can advance as soon as all required children become terminal. Typed parent scoped failures suppress undispatched siblings of the failed parent while allowing unrelated parents to continue. On NVIDIA H20 GPUs, RayOrch achieves 15.14 times speedup when scaling MinerU from 4 to 64 GPUs and 7.82 times speedup when scaling a video pipeline from 8 to 64 GPUs. It reduces end to end time by 13.1 percent versus Ray Data and 29.0 percent versus Daft on MinerU, and by 16.0 percent versus Ray Data on Docling. Code available at https://github.com/OpenDCAI/RayOrch .
View arXiv page
View PDF
Project page
GitHub
16
Add to collection
Community
Sunnyhaze
Paper author
Paper submitter
about 16 hours ago
We introduce RayOrch, a distributed data pipeline system for multimodal foundation model data preparation. Like Ray Data and Daft, it targets large scale data processing. RayOrch is designed for pipelines with multiple stages and levels of granularity, such as processing documents as pages and videos as frames. It tracks which input each piece belongs to, batches ready work across inputs for GPU execution, and gathers results in the correct order even when tasks finish out of order. On the MinerU workload, it reduces end to end time by 13.1% compared with Ray Data and 29.0% compared with Daft. We welcome feedback and questions!
See translation
🔥
1
+
Reply
researchstudio-bot
about 13 hours ago
This is an automated message from the ResearchStudio team.
We created an interactive ResearchStudio Reel for this paper. It includes a visual poster, a video, and a blog, all available for download in editable formats.
Open the ResearchStudio Reel →
Download all files from Hugging Face
Please give this comment a thumbs up if you find the Reel helpful!
Want to explore or create Reels for more papers? Visit the ResearchStudio demo.
See translation
Reply
Edit
Preview
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Comment
· Sign up or log in to comment
Upvote
43
+31
Get this paper in your agent:
hf papers read 2609.18703
Don't have the latest CLI?
Models citing this paper
0
No model linking this paper
Cite arxiv.org/abs/2609.18703 in a model README.md to link it from this page.
Datasets citing this paper
0
No dataset linking this paper
Cite arxiv.org/abs/2609.18703 in a dataset README.md to link it from this page.
Spaces citing this paper
0
No Space linking this paper
Cite arxiv.org/abs/2609.18703 in a Space README.md to link it from this page.
Collections including this paper
3
Code
Collection
82 items
•
Updated about 11 hours ago
•
7
Dataset and Data processing
Collection
10 items
•
Updated about 11 hours ago
1
Collection
6 items
•
Updated about 3 hours ago
首次收录 · 2026-09-29 · 10.95 分