What changed
Three new datasets have been added to Hugging Face, all contributed by the user 'bermaneh'. These datasets are designed for the evaluation of Large Language Models (LLMs) in the context of Partial Differential Equations (PDEs), specifically focusing on free-generation tasks. The datasets include:
* `bermaneh/pde-llm-eval-freegen-xmodal-qwen__qwq-32b`: This dataset contains 768 rows and 37 columns, with both tabular and text modalities. It is licensed under MIT and formatted in Parquet. * `bermaneh/pde-llm-eval-freegen-xmodal-qwen__qwen3-32b`: Similar to the first, this dataset also features 768 rows and 37 columns, supporting tabular and text data, with an MIT license and Parquet format. * `bermaneh/pde-llm-eval-freegen-xmodal-deepseek-ai__deepseek-r1-distill-qwen-32b`: This dataset also comprises 768 rows and 37 columns, suitable for tabular and text data, and is released under the MIT license in Parquet format.
All datasets are marked as 'final' and are intended for free-generation evaluations.
Why it matters for builders
For AI builders and researchers, these datasets offer a standardized way to benchmark and compare the performance of different LLMs on complex scientific problems like PDEs. This is crucial for developing more capable AI models that can assist in scientific discovery and engineering simulations. The inclusion of both tabular and text data allows for testing multimodal capabilities.
Practical impact
Developers can leverage these datasets to fine-tune their models for scientific reasoning or to evaluate how well existing models generalize to PDE-related tasks. The availability of these specific evaluation sets can accelerate the development cycle for domain-specific LLMs, leading to more accurate and reliable AI tools for scientific communities.
Caveats and source limits
The provided sources are Hugging Face dataset signals and do not contain detailed benchmark results or performance comparisons between models. The datasets themselves are described as having 768 rows and 37 columns, with a focus on free-generation evaluation for PDEs. No specific model names beyond those used in the dataset IDs (e.g., Qwen, DeepSeek) are explicitly evaluated or compared within the provided excerpts. The 'likes' and 'downloads' counts for these datasets are currently zero, indicating they are newly added or have not yet gained traction.
Featured on AI Radar: New Hugging Face Datasets for PDE LLM Evaluation