Data Summarization at Scale: A Two-Stage Submodular Approach

Marko Mitrovic; Ehsan Kazemi; Morteza Zadimoghaddam; Amin Karbasi

首页> 外文期刊>JMLR: Workshop and Conference Proceedings >Data Summarization at Scale: A Two-Stage Submodular Approach

【24h】

Data Summarization at Scale: A Two-Stage Submodular Approach

机译：大规模数据汇总：两阶段亚模方法

获取原文

掌桥外文数据库（机构版） >>

开具论文收录证明 >>

文献代查 >>

页面导航

摘要
著录项
相似文献
相关主题

摘要

The sheer scale of modern datasets has resulted in a dire need for summarization techniques that can identify representative elements in a dataset. Fortunately, the vast majority of data summarization tasks satisfy an intuitive diminishing returns condition known as submodularity, which allows us to find nearly-optimal solutions in linear time. We focus on a two-stage submodular framework where the goal is to use some given training functions to reduce the ground set so that optimizing new functions (drawn from the same distribution) over the reduced set provides almost as much value as optimizing them over the entire ground set. In this paper, we develop the first streaming and distributed solutions to this problem. In addition to providing strong theoretical guarantees, we demonstrate both the utility and efficiency of our algorithms on real-world tasks including image summarization and ride-share optimization.

机译：现代数据集的庞大规模导致迫切需要能够识别数据集中代表性元素的汇总技术。幸运的是，绝大多数数据汇总任务满足了称为子模块化的直观递减条件，这使我们能够在线性时间内找到近乎最优的解决方案。我们关注于一个两阶段的子模块框架，其目标是使用一些给定的训练函数来减少地面集合，从而在减少的集合上优化新函数（从相同分布中提取）提供的价值几乎与在地面上优化它们的价值相同。整个地面。在本文中，我们开发了第一个针对此问题的流式和分布式解决方案。除了提供强有力的理论保证外，我们还演示了算法在实际任务中的效用和效率，包括图像摘要和乘车共享优化。

著录项

来源
《JMLR: Workshop and Conference Proceedings》 |2018年第1期|共10页
作者
Marko Mitrovic; Ehsan Kazemi; Morteza Zadimoghaddam; Amin Karbasi;
展开▼
作者单位

展开▼
收录信息
原文格式 PDF
正文语种
中图分类人工智能理论;
关键词

相似文献

外文文献
中文文献
专利

1. Data Summarization at Scale: A Two-Stage Submodular Approach [J] . Marko Mitrovic, Ehsan Kazemi, Morteza Zadimoghaddam, JMLR: Workshop and Conference Proceedings . 2018,第12期

机译：大规模数据汇总：两阶段亚模方法
2. Scalable Deletion-Robust Submodular Maximization: Data Summarization with Privacy and Fairness Constraints [J] . Ehsan Kazemi, Morteza Zadimoghaddam, Amin Karbasi JMLR: Workshop and Conference Proceedings . 2018,第2010期

机译：可扩展的删除 - 强大的子模块化最大化：数据摘要与隐私和公平约束
3. Approximation Algorithms for Submodular Data Summarization with a Knapsack Constraint [J] . Kai Han, Enpei Zhang, Tong Xu, Performance evaluation review . 2021,第1期

机译：带有背包约束的子模块数据摘要的近似算法
4. Differentially Private Submodular Maximization: Data Summarization in Disguise (Full version) [C] . Marko Mitrovic, Mark Bun, Andreas Krause, International Conference on Machine Learning . 2018

机译：差异私有子模块最大化：伪装数据摘要（完整版）
5. In Situ Summarization and Visual Exploration of Large-Scale Simulation Data Sets [D] . Dutta, Soumya. 2018

机译：大规模仿真数据集的原位汇总和可视化探索
6. A summarization approach for Affymetrix GeneChip data using a reference training set from a large biologically diverse database [O] . Simon Katz, Rafael A Irizarry, Xue Lin, 2006

机译：Affymetrix GeneChip数据的汇总方法使用来自大型生物多样性数据库的参考训练集
7. Approximation Algorithms for Submodular Data Summarization with a Knapsack Constraint [O] . Kai Han, Shuang Cui, Tianshuai Zhu, 2021

机译：近似算法与背包约束的子模块数据摘要

Data Summarization at Scale: A Two-Stage Submodular Approach

摘要

著录项

相似文献

相关主题

期刊订阅