Star 历史趋势
数据来源: GitHub API · 生成自 Stargazers.cn
README.md

Diwan (ديوان): The Largest Classified Dataset of Arabic Poetry

⚠️ Notice:
Dataset files are temporarily unavailable during repository maintenance.

Diwan is the largest classified dataset of Arabic poetry, consisting of nearly half a million poems and over 15 million individual verses. This comprehensive collection covers a wide range of poetic forms, meters, styles, and themes from the beginning of Arabic poetry to the present day. The dataset is structured into various categories to facilitate research in Arabic literature, prosody, and natural language processing.


Table of Contents


Project Overview

Diwan (ديوان) aims to archive and classify Arabic poetry, serving as a valuable resource for researchers, linguists, and enthusiasts. With data collected from various sources, including web archives, poet websites, and historical books, Diwan offers a structured and searchable dataset that spans multiple eras, countries, and poetic forms.

This project was developed with the following goals:

  1. To cover all poems from the beginning of Arabic poetry to the modern era.
  2. To classify these poems into detailed categories based on poetic genre, style, language, meter, period, country, and other characteristics.

The dataset can be used for a variety of tasks, including meter identification, automatic poem generation, topic classification, plagiarism detection, and more.


Current Release Information

  • Total Poems: Nearly 400,000 poems
  • Total Verses: Over 14 million verses

The dataset is categorized by poetic genre, meter, style, theme, and other features, making it suitable for a wide range of applications in research and development.


Explore Diwan on Google Colab

To explore the full details of the Diwan dataset and interact with its extensive features, access our interactive Google Colab notebook. This notebook provides an easy-to-use interface for searching, analyzing, and learning more about the Diwan dataset.

Open in Colab

Steps to Use Diwan on Google Colab:

  1. Run the first cell by clicking the "Run" button to load the necessary files and libraries.
  2. Execute the notebook, and dropdown menus will appear with various poetry categories. You can experiment by selecting different options, and the corresponding poetic data will be displayed based on your choices.
  3. You can explore additional features of the Diwan dataset by accessing the main menu on the left side of the notebook interface.

Future Updates

We are currently in the process of increasing the dataset to:

  • Total Poems: Nearly 500,000 poems
  • Total Verses: Over 15 million individual verses

Additional classifications and refinements will be included in upcoming releases to further enhance the richness of the Diwan dataset.


How to Use the Dataset

The Diwan dataset can be used for various research purposes:

  • Meter identification: Detect the prosodic meter of Arabic poems.
  • Topic classification: Automatically classify poems by theme or purpose.
  • Poem generation: Train models to generate Arabic poetry.
  • Plagiarism detection: Use the dataset to identify duplicate or plagiarized content.
  • Multi-label classification: Explore how poems can belong to multiple categories simultaneously.

License

This project is licensed under the MIT License. See the LICENSE file for more details.


Contributing

We welcome contributions from the community. Please feel free to open an issue or submit a pull request if you have any suggestions or improvements.


关于 About

Diwan is the largest Arabic poetry dataset, containing nearly 500,000 poems and over 15 million verses. It spans various poetic forms, meters, styles, and themes, from ancient to modern times. The dataset is organized into categories to support research in Arabic literature, prosody, and natural language processing.

语言 Languages

Python100.0%

提交活跃度 Commit Activity

代码提交热力图
过去 52 周的开发活跃度
3
Total Commits
峰值: 3次/周
Less
More

核心贡献者 Contributors