RSS Amplifier

Blog

Academic Torrents

Recent Torrents

academictorrents.comSource feed ↗31 posts

Overdue Last read · next check
Last read 4 days ago, longer than this feed's 1 day schedule.

Latest posts

Stack Exchange Data Dump (2026-06-30)

This data dump is sourced from the various sites in the Stack Exchange network of Q&A sites. This dump contains data up to and including 2026-06-30. The exact licenses for each bit of content is embedded in each entry. For license date ranges, see the root-level license.txt, or https://stackoverflow.com/help/licensing. For the schema, see the sede-and-data-dump-schema.md file within each .7z This…

Reddit comments/submissions 2005-06 to 2024-12

Reddit comments and submissions from 2005-06 to 2024-12 collected by pushshift and u/RaiderBDev. These are zstandard compressed ndjson files. Example python scripts for parsing the data can be found here https://github.com/Watchful1/PushshiftDumps The more recent dumps are collected by u/RaiderBDev License: No license specified, the work may be protected by copyright.

Reddit comments/submissions 2026-07

Reddit comments and submisReddit comments and submissions from 2026-07 Documentation, json schemas and more can be found at https://github.com/ArthurHeitmann/arctic_shift Helper scripts for processing files can be found at https://github.com/Watchful1/PushshiftDumpssions

What Should I Become? When LLMs Present a Slice of Opportunity as the Whole

Large language models (LLMs) are increasingly consulted for life-path guidance, yet the distribution of options they surface has gone largely unexamined. We characterize occupational-category visibility across 165,000 LLM responses spanning 100 user profiles, 15 prompts (grouped into seven framings), 11 models across seven families, two role framings, and five temperatures, with keywords derived…

Reddit comments/submissions 2026-06

Reddit comments and submisReddit comments and submissions from 2026-06 Documentation, json schemas and more can be found at https://github.com/ArthurHeitmann/arctic_shift Helper scripts for processing files can be found at https://github.com/Watchful1/PushshiftDumpssions

LUMINOUS Database: Lumbar Multifidus Muscle Segmentation From Ultrasound

This database provides the US ground truth of the left and right LM muscles at the L5 level (in prone and standing positions) of 109 US datasets of young athletic adult volunteers (64 males, 45 females, age: 21.1 ± 1.7). The LUMINOUS database contains the US images with their corresponding manually segmented binary masks, serving as the ground truth. The purpose of the database is to enable…

FALLMUD : FAscicle Lower Leg Muscle Ultrasound Dataset

FAscicle Lower Leg Muscle Ultrasound Dataset is a dataset composed of 812 ultrasound images of lower leg muscles to analyze muscle weaknesses and prevent injuries. This dataset is presented in the article AW-Net: Automatic muscle structure analysis on B-mode ultrasound images for injury prevention. It combines the datasets provided by two articles, “Estimating Full Regional Skeletal Muscle Fibre…

Wikipedia European languages 2026-06-01

Wikipedia database dumps of European language wikis of 10k articles or more. enwiki excluded. Wikipedia Multistream 2026-06-01. These 67 languages are included: Albanian, Alemannic, Aragonese, Asturian, Basque, Bavarian, Belarusian, Benetian, Bosnian, Breton, Bulgarian, Catalan, Croatian, Czech, Danish, Dutch, Emilian-Romagnol, Esperanto, Estonian, Faroese, Finnish, French, Galician, German,…

Reddit comments/submissions 2026-05

Reddit comments and submisReddit comments and submissions from 2026-05 Documentation, json schemas and more can be found at https://github.com/ArthurHeitmann/arctic_shift Helper scripts for processing files can be found at https://github.com/Watchful1/PushshiftDumpssions

Crossref Event Data Archive

# Crossref Event Data archive This is an archive of all events collected by selected Crossref Event Data agents between its launch on 2017/02/17 and its deprecation on 2026/04/23. The DOI of this dataset is https://doi.org/10.13003/wjyr-rv9j ## File format The data are provided in [JSONL](https://jsonlines.org/) format. Each data file has a .jsonl file extension and contains up to 5000 entries…

GTDB R09-RS220

Release 09-RS220 (24th April 2024) of the Genome Taxonomy Database (GTDB), an initiative to establish a standardised microbial taxonomy based on genome phylogeny.

GTDB R07-RS207

Release 07-RS207 (8th April 2022) of the Genome Taxonomy Database (GTDB), an initiative to establish a standardised microbial taxonomy based on genome phylogeny.

GTDB R06-RS202

Release 06-RS202 (27th April 2021) of the Genome Taxonomy Database (GTDB), an initiative to establish a standardised microbial taxonomy based on genome phylogeny.

GTDB R04-RS89

Release 04-RS89 (19th June 2019) of the Genome Taxonomy Database (GTDB), an initiative to establish a standardised microbial taxonomy based on genome phylogeny.

GTDB R03-RS86.2

Release 3-RS86.2 (15th January 2019) of the Genome Taxonomy Database (GTDB), an initiative to establish a standardised microbial taxonomy based on genome phylogeny.

GTDB R05-RS95

Release 05-RS95 (17th July 2020) of the Genome Taxonomy Database (GTDB), an initiative to establish a standardised microbial taxonomy based on genome phylogeny.

GTDB R03-RS86

Release 3-RS86 (19th August 2018) of the Genome Taxonomy Database (GTDB), an initiative to establish a standardised microbial taxonomy based on genome phylogeny.

GTDB R01-RS80

Release 1-RS80 (1st November 2017) of the Genome Taxonomy Database (GTDB), an initiative to establish a standardised microbial taxonomy based on genome phylogeny.

GTDB R02-RS83

Release 2-RS83 (8th March 2018) of the Genome Taxonomy Database (GTDB), an initiative to establish a standardised microbial taxonomy based on genome phylogeny.

GTDB R11-RS232

Release 11-RS232 (15th April 2026) of the Genome Taxonomy Database (GTDB), an initiative to establish a standardised microbial taxonomy based on genome phylogeny.

enwiki-20260601-pages-articles-multistream.xml.bz2

English Wikipedia Multistream 2026-06-01 https://en.wikipedia.org/wiki/Wikipedia:Database_download Corresponding index file: https://academictorrents.com/details/c1236c4d35b6d2adcba502e3271d6a3c5261b1ab

enwiki-20260601-pages-articles-multistream-index.txt.bz2

English Wikipedia Multistream Index 2026-06-01 https://en.wikipedia.org/wiki/Wikipedia:Database_download Corresponding multistream file: https://academictorrents.com/details/bac5df1f39fd83fc87826a8dc546e56db34f2322

Places in the Wild: Ecologically-sampled RAW photographs

Places in the Wild comprises over 67,000 RAW-format images, each captured with a 45-megapixel Canon EOS R5 full-frame mirrorless camera at 5-degree intervals, providing 360-degree coverage across over 800 unique locations. These locations span 260 basic-level scene categories, including both indoor and outdoor environments such as bedrooms, train stations, forests, and parking garages.

Edus2 Ultrasounds

Ultrasound Videos Database: Collection of 32 medical ultrasound video files for simulations, case discussions, and training. Includes cardiac normal, tamponade, FAST exams (RUQ free fluid), AAA, and Edus2 open-source set. Free for non-commercial educational use. This license applies to all video in this directory. Copyright 2011,2012 Paul Kulyk and Paul Olszynski All videos made available under a…

Reddit comments/submissions 2026-04

Reddit comments and submisReddit comments and submissions from 2026-04 Documentation, json schemas and more can be found at https://github.com/ArthurHeitmann/arctic_shift Helper scripts for processing files can be found at https://github.com/Watchful1/PushshiftDumpssions

Wikipedia Asian languages 2026-05-01

Wikipedia database dumps of Asian language wikis of 10k articles or more. Wikipedia Multistream 2026-05-01. These 85 languages are included: Acehnese, Armenian, Assamese, Azerbaijani, Balinese, Bangla, Banjar, Banyumasan, Bashkir, Bishnupriya, Buginese, Burmese, Cantonese, Cebuno, Central Bikol, Central Kurdhish, Chechen, Chinese, Chuvash, Classical Chinese, Dimli, Eastern Mari, Georgian, Gilaki,…

Stack Exchange Data Dump (2026-03-31)

This data dump is sourced from the various sites in the Stack Exchange network of Q&A sites. This dump contains data up to and including 2026-03-31. The exact licenses for each bit of content is embedded in each entry. For license date ranges, see the root-level license.txt, or https://stackoverflow.com/help/licensing. For the schema, see the sede-and-data-dump-schema.md file within each .7z This…

IUGC: A benchmark of landmark detection in end-to-end intrapartum ultrasound biometry

In 2018, the World Health Organization (WHO) published 56 recommendations to improve the quality of intrapartum care and enhance women’s childbirth experiences. In response, the WHO developed the Labour Care Guide (LCG) in 2020, a next-generation tool designed to promote evidence-based, respectful, and woman-centered care during labor and delivery. The LCG was created through expert consultations,…

Annotated Ultrasound Liver images

We public the ultrasound liver images, which were annotated to show the outlines, livers, and liver mass regions. Xu Yiming, Zheng Bowen, Liu Xiaohong, Wu Tao, Ju Jinxiu, Wang Shijie, Lian Yufan, Zhang Hongjun, Liang Tong, Sang Ye, Jiang Rui, Wang Guangyu, Ren Jie, & Chen Ting. (2022). Annotated Ultrasound Liver images [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7272660

TRUSTED: The Paired 3D Ultrasound and CT Human Data for Kidney Segmentation and Registration Research

We propose TRUSTED (the Tridimensional Renal Ultra Sound Tomod Ensitometrie Dataset), comprising paired transabdominal 3DUS and CT kidney images from 48 human patients (96 kidneys), including segmentation, and anatomical landmark annotations by two experienced radiographers. Abstract Inter-modal image registration (IMIR) and image segmentation with abdominal Ultrasound (US) data have many…

Abdominal Ultrasound Image Dataset for Organ Classification and Disease Detection

This is a dataset of Ultrasound (US) images of abdominal organs. US imaging is widely accessible and a very common diagnostic tool, as it is non-invasive and does not involve radiation risk. This dataset was curated solely for research in deep learning, with potential applications in supervised, semi-supervised, and unsupervised learning to support disease detection in resource-constrained…