新闻标题分类数据集(HuffPost)
HuffPost 新闻标题与类目(约 20 万条,2012-2022),文本分类经典数据集。

数据预览 · 记录样例
数据开头的 5 条真实新闻记录
Health experts said it is too early to predict whether demand would match up with the 171 million doses of the new boosters the U.S. ordered for the fall.
2022-09-23He was subdued by passengers and crew when he fled to the back of the aircraft after the confrontation, according to the U.S. attorney's office in Los Angeles.
2022-09-23"Until you have a dog you don't understand what could be eaten."
2022-09-23"Accidentally put grown-up toothpaste on my toddler’s toothbrush and he screamed like I was cleaning his teeth with a Carolina Reaper dipped in Tabasco sauce."
2022-09-23Amy Cooper accused investment firm Franklin Templeton of unfairly firing her and branding her a racist after video of the Central Park encounter went viral.
2022-09-22数据交付
本条目提供数据和数据说明,暂未提供可运行教学案例。需要自行完成建模或实验。
数据集介绍
【数据集名称】新闻标题分类数据集(HuffPost) 【门类归属】艺术与文化 / 语言数据 【数据类型与任务】文本;文本分类(新闻类目) 【数据规模】209527 条记录 【文件格式】JSON 文件 【压缩包大小】下载压缩包约 26.5 MB 【数据划分】资料合集类数据,非模型训练用途,无训练集/验证集/测试集划分。 【标签或字段】有标签:category 字段,42 个类目;完整的字段与类别说明见压缩包内的数据字典文件 data-dictionary.csv。 【类型关键参数】共 209527 条记录,JSON Lines 格式(一个文件里一行一条记录)。主要字段:link(链接)、headline(标题)、category(类别)、short_description(摘要)、authors(作者)、date(日期)。内容为英文。 【目录结构】 news-category-dataset.zip(解压后) News_Category_Dataset_v3.json 【怎么使用】 整个数据集是一个 JSON 文本文件,一行一条记录,用记事本都能打开看;用 Python 的 pandas 一行代码就能读成表格: import pandas as pd df = pd.read_json("News_Category_Dataset_v3.json", lines=True) # 209527 条记录 每个字段的含义见压缩包内的 data-dictionary.csv。 【数据体检结果】 · 数据体检未发现问题,可放心使用。
整理与使用信息
本站处理内容
本站已完成的工作: 1. 中文名称与用途说明整理(对照原站信息逐项核对); 2. 全量文件清点与 SHA256 校验(file-manifest.csv,可溯源); 3. 自动质检报告(quality-report.json:结构、编码、重复、标注一致性等); 4. 数据字典整理(data-dictionary.csv,字段/类别中英对照); 5. 预览样例抽取(preview/,供购买前判断内容)。
中文整理说明
门类:艺术与文化 / 语言数据(依据《数据集中文分类体系》) 中文命名:新闻标题分类数据集(HuffPost)(原始名称 rmisra/news-category-dataset) 已知问题:自动检查未发现问题(详见 quality-report.json) 配套文档:简明说明.md、README_中文.md、数据字典、文件清单、质检报告、预览样例(均随网盘下载包提供)。