Credits & Licences
TutorLens is built on a small number of openly licensed datasets. Several of those licences ask us to credit their authors somewhere you can actually reach — this is that page. Questions: privacy@tutorlens.co.uk.
Word frequency
SUBTLEX-UK— van Heuven, W.J.B., Mandera, P., Keuleers, E., & Brysbaert, M. (2014). SUBTLEX-UK: A new and improved word frequency database for British English. Quarterly Journal of Experimental Psychology, 67(6), 1176–1190. Zipf frequencies are used under the SUBTLEX-UK terms (commercial use permitted with attribution). SUBTLEX-UK is free to obtain from its authors.
Lexical databases
WordNet — Princeton University. WordNet: A Lexical Database for English. Used under the WordNet Licence.
Open English WordNet 2024 — licensed under CC BY 4.0.
Roget's Thesaurus (1911)— public domain. Peter Mark Roget's thesaurus, 1911 edition, sourced from Project Gutenberg; used to find candidate pairs of opposites. No Roget text is shown to you.
Coxhead Academic Word List — Coxhead, A. (2000). A New Academic Word List. TESOL Quarterly, 34(2), 213–238.
Concreteness norms— Brysbaert, M., Warriner, A.B., & Kuperman, V. (2014). Concreteness ratings for 40 thousand generally known English word lemmas. Behavior Research Methods, 46, 904–911. Used inside our own tooling only, to decide how hard a word is — nothing from it is shown to you.
Word lists and filters
SCOWL — Spell Checker Oriented Word Lists (SCOWL) by Kevin Atkinson. Used under the SCOWL licence, to check that a word is valid British English.
LDNOOBW — List of Dirty, Naughty, Obscene, and Otherwise Bad Words. Licensed under CC BY 4.0. Used to keep unsuitable words out of questions written for children.
Curriculum
Contains public sector information licensed under the Open Government Licence v3.0. Source: National Curriculum in England, English Appendix 1 (spelling), Department for Education.
Reading passages
Some comprehension passages are public-domain works sourced from Project Gutenberg (public domain). Project Gutenberg trademark lines and licence headers are stripped before use; only the public-domain text itself is shown. Other passages are written for us and are original work.
A separate set of public-domain Project Gutenberg children's books is read only to measure how common and how difficult words are. Nothing from those books is stored or shown — just the numbers we work out from them.
Everything else
Questions, lessons, explanations and Arlo's hints are our own work. Wikidata is used under CC0 and needs no attribution; we credit it here for completeness.
Full provenance, including the sources we deliberately chose not to use, is recorded internally in our data licence register.