소비자Synthetic data

Women's E Commerce Clothing Reviews Dataset

Main product

Data Quantity

Optional product

During synthetic data generation, duplication may occur due to the close resemblance to the original dataset, a common issue in such processes. To minimize this, consider generating more data than initially required.

Total Price

USD 0 VAT included

  • Main product
    Premium
  • Data Quantity
    Basic
  • Optional product
    Not selected
  • Download
    No data
  • All products are priced including VAT.
  • Premium datasets are custom-made and take approximately two weeks from the date of application to improve quality.
  • Common datasets can be checked on My Page after purchase.
Product Image
Consumer

Basic Price

USD 4,400 VAT included

  • Labeling type: Rating
  • Data Format: Text

Related Tags:

Use cases of research datasets
No thing

About Dataset

1) Data introduction

• Womens-ecommerce-clothing-reviews dataset is a dataset containing 23,000 customer reviews and ratings.

2) Data utilization

(1) Womens-ecommerce-clothing-reviews data has characteristics that: • We aim for high-quality NLP and multivariate analysis with a dataset consisting of 10 functional variables such as clothing, age, and review title and 23,486 rows. (2) Womens-ecommerce-clothing-reviews data can be used to: • Rating prediction: Develop machine learning models to predict the ratings customers might give based on review text and support automated review analysis. • Trend analysis: Companies can analyze data to identify trends and patterns in customer preferences and support inventory management and marketing strategies.

Meta Data

DomainConsumerZoodata formatsText
Zoodata volume1000 itemsRegistration date2025.02.26
Zoodata typeSynthetic dataExistence of labelingExist
Labeling typeRatingLabeling formatsjson

Good

Performance 1
85

Outstanding

Performance 2
100

Data Samples

Sample Data

Utility

Downstream Classification (▲)MMD (▼)One Class Classification (▼)
Total000
SuitabilityOKOKOK

The higher the value, the better (▲)

Model Performance

Downstream classification accuracy is an indicator used to evaluate the usefulness of synthetic data. It measures whether synthetic data performs similarly to real data. The method involves training the same model separately on real data and synthetic data, and then comparing the accuracies of the two models. Interpretation: A high accuracy rate means that the model trained on synthetic data performs similarly to the one trained on real data, indicating that the synthetic data is of high quality and well represents the real data.

The closer to zero or the lower the value, the better (▼)

Quality

MMD (Maximum Mean Discrepancy) is a metric used to assess the similarity between two probability distributions. It is commonly used to compare generated data with real data. High MMD score: A score above 0.05 indicates that the two distributions may differ. Low MMD score: Indicates that the generated data is similar to the real data. A score close to 0 is preferable, and a score below 0.01 suggests that the two data distributions are nearly indistinguishable.


Privacy

ROUGE (▼)BERTScore (▼)
Total00
SuitabilityOKOK

The closer to zero or the lower the value, the better (▼)

Structural Similarity

Inference risk measures the risk of inferring sensitive information about the original data from synthetic data. It evaluates the likelihood of extracting original data information from synthetic data and is calculated by comparing the distance between synthetic and original data. High Duplication Rate: Indicates lower data diversity and potential quality issues, which can reduce the reliability of analysis and models. Interpretation: A lower risk value means that synthetic data is less likely to infer sensitive information, indicating higher data security.

Perceptual Similarity

BERTScore is a metric used to evaluate the semantic similarity between two texts, utilizing deep learning models like BERT (Bidirectional Encoder Representations from Transformers) to capture the meaning of the texts.. High BERTScore: If the value is 0.9 or higher, it indicates that the generated text is almost semantically identical to the real text. This can increase the risk of sensitive information leakage. Low BERTScore: If the value is 0.6 or lower, it indicates lower similarity and a reduced risk of leakage.

Premium Report Information

If you purchase the premium report product, you will be able to view the analysis results of a more detailed dataset.
select premium data

Premium dataset sample