Harry Potter Reviews Dataset

A comprehensive analysis of 491 reviews of the famous "Harry Potter and the Sorcerer's Stone," which can be used for sentiment analysis, marketing and sales analytics, and other data-driven applications.

Related Tags

Total Price

$ 2,500

(VAT Included)

Looking for custom-made dataset or researcher-accessible data? Please contact us for inquiries.

About Dataset

1) Data introduction

• Harry-potter-reviews dataset contains 491 comprehensive reviews of the famous “Harry Potter and the Sorcerer’s Stone.”

2) Data utilization

(1)Harry-potter-reviews data has characteristics that: • Contains 491 comprehensive reviews of the film. The reviews were generated using Large Language Models (LLMs). (2) Harry-potter-reviews data can be used to: • Sentiment analysis: This dataset can be used to train a model that measures the overall reader satisfaction by analyzing the sentiment of reviews and classifying them as positive, negative, or neutral. • Marketing and sales analytics: Publishers and marketers can use insights from data to understand reader preferences and improve marketing strategies for similar books.

Meta Data

Domain	Consumer	Zoodata formats	Text
Zoodata volume	1000 items	Registration Date	2025.02.26
Zoodata type	Synthetic data	Existence of labeling	Exist
Labeling type	Rating	Labeling formats	json

Good

Outstanding

100

Data Samples

Utility

	Downstream Classification (▲)	MMD (▼)	One Class Classification (▼)
Total	0	0	0
Suitability	OK	OK	OK

The higher the value, the better (▲)

Model Performance

Downstream classification accuracy is an indicator used to evaluate the usefulness of synthetic data. It measures whether synthetic data performs similarly to real data. The method involves training the same model separately on real data and synthetic data, and then comparing the accuracies of the two models. Interpretation: A high accuracy rate means that the model trained on synthetic data performs similarly to the one trained on real data, indicating that the synthetic data is of high quality and well represents the real data.

The closer to zero or the lower the value, the better (▼)

Quality

MMD (Maximum Mean Discrepancy) is a metric used to assess the similarity between two probability distributions. It is commonly used to compare generated data with real data. High MMD score: A score above 0.05 indicates that the two distributions may differ. Low MMD score: Indicates that the generated data is similar to the real data. A score close to 0 is preferable, and a score below 0.01 suggests that the two data distributions are nearly indistinguishable.

Privacy

	ROUGE (▼)	BERTScore (▼)
Total	0	0
Suitability	OK	OK

The closer to zero or the lower the value, the better (▼)

Structural Similarity

Inference risk measures the risk of inferring sensitive information about the original data from synthetic data. It evaluates the likelihood of extracting original data information from synthetic data and is calculated by comparing the distance between synthetic and original data. High Duplication Rate: Indicates lower data diversity and potential quality issues, which can reduce the reliability of analysis and models. Interpretation: A lower risk value means that synthetic data is less likely to infer sensitive information, indicating higher data security.

Perceptual Similarity

BERTScore is a metric used to evaluate the semantic similarity between two texts, utilizing deep learning models like BERT (Bidirectional Encoder Representations from Transformers) to capture the meaning of the texts.. High BERTScore: If the value is 0.9 or higher, it indicates that the generated text is almost semantically identical to the real text. This can increase the risk of sensitive information leakage. Low BERTScore: If the value is 0.6 or lower, it indicates lower similarity and a reduced risk of leakage.

Premium Report Information

If you purchase the premium report product, you will be able to view the analysis results of a more detailed dataset.

Premium dataset sample