Back to Bingelang
Sources and licenses

Open lexical data, with its credits attached

Bingelang builds its lexical index from documented frequency data and language models. This page identifies that layer, its sources, and the licenses that travel with it.

Last updated: July 20, 2026

What the lexical layer contains

The separable lexical layer contains canonical lemmas, parts of speech, normalized surface forms, source-derived frequency scores, and ranks. It does not contain Bingelang-authored definitions, translations, examples, audio, interface code, or user data.

French is the first candidate language. It will remain unpublished until its ambiguity, display-form, and manual-review quality gates pass.

wordfreq 3.1.1

wordfreq by Robyn Speer and contributors supplies the source frequency probabilities. Its software is licensed under Apache 2.0. The project states that its redistributable data files are available under Creative Commons Attribution-ShareAlike 4.0.

The combined data credits include Wikipedia, OPUS OpenSubtitles, SUBTLEX authors, Google Books Ngrams, Leeds Internet Corpus, ParaCrawl, GlobalVoices, News Crawl, OSCAR, Reddit, and aggregate Twitter statistics. Full source-specific conditions remain in the upstream license notice.

The Bingelang lexical layer derived from this data is distributed under CC BY-SA 4.0 with its attribution and release metadata. It is not distributed as a detached CSV that could lose those notices.

Stanza 1.14.0

Stanza by the Stanford NLP Group supplies tokenization, multi-word-token analysis, morphology, universal part-of-speech tags, and lemmas. The software is licensed under Apache 2.0.

Stanford describes its language packs, to the extent Stanford owns the relevant rights, under the Open Data Commons Attribution License 1.0. Each Bingelang release records the exact processors, packages, model files, and hashes used.

Citation: Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton and Christopher D. Manning. 2020. Stanza: A Python Natural Language Processing Toolkit for Many Human Languages.

French GSD model

The French candidate uses Stanza's fr_gsd package, trained from the Universal Dependencies French GSD treebank. The treebank is licensed under CC BY-SA 4.0.

The repository credits Marie-Catherine de Marneffe, Bruno Guillaume, Matias Grioni, Carly Dickerson, Guy Perrier, Ryan McDonald, Alane Suhr, Joakim Nivre, and additional contributors.

Lexique 4.00

Lexique 4.00 independently checks French inflected forms, lemmas, grammatical categories, gender, and number. It is distributed under CC BY-SA 4.0. Lexique supplies no Bingelang frequency score and cannot reorder the wordfreq ranking.

Citation: Boris New, Christophe Pallier, Claudia Schalchli, Jessica Bourgin, and Jean-Baptiste Gimenes. 2026. Lexique 4: A major upgrade of the Lexique French lexical database.

How attribution stays attached

Source versions, license metadata, model provenance, and the generation manifest are stored with each lexical release. Future languages will not be published until their own model and corpus notices have been added.

Questions about these notices can be sent to contact@bingelang.com.