WangNet

A character-level classifier that decides whether a number is Numberwang.

Give it a number, an arithmetic expression, or a quantity in context, in any of eleven languages, and it returns one of four verdicts. 356,868 parameters, trained from scratch on 187,606 labelled examples.

โ–ถ Try it in your browser โ€” runs client-side, no server, no wait.

from transformers import pipeline

wang = pipeline("text-classification", model="graafhenk/numberwang")
wang("22")
# [{'label': 'numberwang', 'score': 0.976}]

Accuracy compared with baselines and a 1.7B language model

The four verdicts

label verdict
not_numberwang That's not Numberwang.
numberwang THAT'S NUMBERWANG!
not_a_number That's not even a number. It can never be Numberwang.
wangernumb That's Wangernumb!

What it accepts

input behaviour
42, sixty-six, 12345 digits or words
zweiundzwanzig, veintidรณs, tweeรซntwintig eleven languages, accents optional
5*2, 96 divided by 2 arithmetic, judged on the result
45 - 44, double four, eins anything worth 1 or 44 is Wangernumb
XLIV, twenty-third, 22nd Roman numerals and ordinals
fortnight, vierendelen, september words built on a number, judged as that number
achtneming, often, money words that merely contain one are not numbers
bonjour, hello how are you no numeric content โ€” can never be Numberwang

A number's wangness is a property of the number, not the language it is written in: four, vier, quatre and cuatro all get the same verdict. Input is lowercased by the tokenizer, so capitalisation makes no difference.

Model details

Architecture BERT, 2 layers, hidden 128, 4 heads
Parameters 356,868
Tokenizer character-level, 69-token vocabulary, lowercasing
Input raw text; the first 64 characters carry the trained signal
Output 4-way softmax
Precision float32

The tokenizer splits on individual characters rather than words, so digits, operators and accented letters are all first-class input and there is no out-of-vocabulary path for ordinary text.

Evaluation

Measured on 486 reserved held-out cases, disjoint from training by construction โ€” roughly 4.3% of every integer range plus a set of phrase templates are reserved before the training data is compiled, so the evaluation set cannot leak as that data grows.

Accuracy 87.4%, macro-F1 0.878.

Precision and recall for each verdict

class precision recall F1 n
Wangernumb 0.941 0.970 0.955 33
Numberwang 0.894 0.913 0.903 241
not Numberwang 0.839 0.818 0.828 165
not a number 0.844 0.809 0.826 47

Confusion matrix over the reserved held-out cases

Errors are concentrated between the two principal verdicts โ€” 43 of the 61 mistakes are Numberwang confused with not Numberwang, and almost all of those are arithmetic the model cannot actually compute. Only four genuine non-numbers were ever called a number.

About 2% of the training labels are deliberately inverted, so no model trained on this data can exceed roughly 98%.

Limitations

  • Arithmetic is memorised, not computed. A classifier cannot perform arithmetic; small operands are covered exhaustively in training and larger ones are guesses. 5*2 is reliable, 904 * 3 is not. If arithmetic correctness matters, evaluate the expression and pass the result.
  • Long input degrades. Only the first 64 characters carry trained signal. Longer text is accepted without error but the tail is not meaningfully read.
  • Verdicts are learned, not guaranteed. There is no rules engine at inference to enforce anything.
  • The ceiling is ~98%, not 100%, by construction.

A zero-dependency alternative

The original WangNet is a hand-rolled convnet stored as a 1.8 MB JSON file, with inference in about 100 lines of pure Python standard library โ€” no torch, no transformers, nothing to install. It scores slightly higher (88.9% on the same held-out set) and the two agree on 88.1% of verdicts, with nearly all disagreements on arithmetic that neither model can actually compute.

That version exists because it deploys anywhere Python does. This one exists because Hugging Face can serve it: github.com/GraafHenk/numberwang.

License

MIT. No warranty is expressed or implied as to whether any particular number is, or is not, Numberwang.

Downloads last month
47
Safetensors
Model size
357k params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using graafhenk/numberwang 1

Evaluation results