Skip to content

Quantitative Analysis of Narrative Reports of Psychedelic Drugs

Jeremy R. Coyle, David E. Presti, Matthew J. Baggott

arXiv Preprint Archive June 1, 2012 via arXiv

Summary

AI-generated from the abstract

Machine learning applied to 1000 written reports of 10 different drugs from the Erowid website identified distinct patterns in how people describe their experiences with each drug. A random-forest classifier using just 110 key words achieved 51.1% accuracy in identifying which drug a report described, far above the 10% expected by chance. Reports of MDMA were most distinctive (86.9% accuracy), while those for DPT were hardest to classify (20.1%). Hierarchical clustering revealed similarities between certain drugs, such as DMT and Salvia divinorum. The findings suggest that automated text analysis can uncover consistent, drug-specific features in subjective experience reports, potentially aiding hypothesis generation about new or poorly understood compounds.

Study at a glance

Characteristics Observational study using machine learning on text reports Peer reviewed
Sample size 1,000
Population Written drug experience reports from the Erowid.org website, covering 10 drugs
Keywords Q-bio.qm Psychedelics Natural language processing Consciousness research
Key finding A random-forest classifier using a subset of 110 words achieved 51.1% accuracy in distinguishing which of 10 drugs a report described, with MDMA reports most accurately classified (86.9%) and DPT reports least accurately (20.1%).

Abstract

Background: Psychedelic drugs facilitate profound changes in consciousness and have potential to provide insights into the nature of human mental processes and their relation to brain physiology. Yet published scientific literature reflects a very limited understanding of the effects of these drugs, especially for newer synthetic compounds. The number of clinical trials and range of drugs formally studied is dwarfed by the number of written descriptions of the many drugs taken by people. Analysis of these descriptions using machine-learning techniques can provide a framework for learning about these drug use experiences. Methods: We collected 1000 reports of 10 drugs from the drug information website Erowid.org and formed a term-document frequency matrix. Using variable selection and a random-forest classifier, we identified a subset of words that differentiated between drugs. Results: A random forest using a subset of 110 predictor variables classified with accuracy comparable to a random forest using the full set of 3934 predictors. Our estimated accuracy was 51.1%, which compares favorably to the 10% expected from chance. Reports of MDMA had the highest accuracy at 86.9%; those describing DPT had the lowest at 20.1%. Hierarchical clustering suggested similarities between certain drugs, such as DMT and Salvia divinorum. Conclusion: Machine-learning techniques can reveal consistencies in descriptions of drug use experiences that vary by drug class. This may be useful for developing hypotheses about the pharmacology and toxicity of new and poorly characterized drugs.

Comments

No comments yet.

Log in to comment