WhichTok? Comparing Three TikTok Data Acquisition Tools
2026-08-10 • Social and Information Networks
Social and Information NetworksHuman-Computer Interaction
AI summaryⓘ
The authors studied how researchers collect data from TikTok, a popular social media platform, and found big differences between three common data collection tools. These tools use different methods, like official APIs or web scraping, which affects the types of posts and time periods they capture. The authors found that only user data was consistent across all tools, while hashtag and keyword data varied a lot. They suggest that these differences might make research findings less reliable and offer advice to improve how TikTok data is collected and reported.
TikToksocial media researchAPIweb scrapingdata collectionalgorithmreproducibilityhashtagskeywordsmethodology
Authors
Gayoung Jeon, Cameron Moy, Silvia Teliz, Cristina Monzer, Nicolette Alayon, Deen Freelon
Abstract
TikTok's global growth has made it a prime platform for both entertainment and political discourse, prompting increased social science research. However, this rapidly evolving research field faces a fundamental reproducibility crisis. TikTok's opaque algorithmic systems hinder researchers from drawing meaningful empirical inferences, while the lack of standardized data collection methods compounds these challenges. This study addresses these methodological gaps by systematically comparing three data collection tools - the official TikTok Research API, Pyktok, and Apify. We evaluated five endpoints: User, Hashtag, Keyword, Comment, and Related Video. Results show substantial cross-tool differences, especially for hashtag and keyword searches. The Research API uses back-end API calls, whereas Apify and Pyktok rely on front-end web scraping, producing systematic differences in the time periods and popularity levels represented in retrieved content. The three tools yielded comprehensive and consistent results only for the user endpoint. Our results question whether these tools can acquire truly random[-ized] samples, as they introduce methodological confounds that may compromise research validity in ways not yet fully understood. Based on these results, we offer methodological, transparency, and ethical recommendations and guidelines to increase TikTok research quality.