A new method using Gaussian Mixture Models and Large Language Models has been proposed to address the issue of imbalanced data clustering in Natural Language Processing, particularly for underrepresented topics in unsupervised tasks. This approach aims to improve the accuracy of clustering algorithms by augmenting targeted data, making it crucial for enhancing the handling of minority topics in NLP.