Instructions to use syvai/danish-pnc with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use syvai/danish-pnc with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="syvai/danish-pnc")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("syvai/danish-pnc") model = AutoModelForTokenClassification.from_pretrained("syvai/danish-pnc", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Danish punctuation restoration and truecasing
This model adds punctuation (, . ? !) and capitalisation to lowercase, unpunctuated Danish text, such as the
output of a speech recogniser. Words are never changed, added or removed, so it can be used on transcripts with word
timestamps. The architecture, base model and label scheme are identical to
RyeAI/ekko-pnc; only the training data differs.
Input: hej mit navn er emil hvordan går det i dag jeg har været på kontoret og bagefter skal jeg hjem
Output: Hej, mit navn er Emil. Hvordan går det i dag? Jeg har været på kontoret, og bagefter skal jeg hjem.
How to use
With transformers (PyTorch)
from transformers import AutoTokenizer, AutoModelForTokenClassification
import torch
repo = "syvai/danish-pnc"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForTokenClassification.from_pretrained(repo).eval()
labels = model.config.id2label
def punctuate(text):
words = text.split()
enc = tok(words, is_split_into_words=True, return_tensors="pt", truncation=True, max_length=256)
with torch.no_grad():
pred = model(**enc).logits[0].argmax(-1).tolist()
out, seen, cap_next = [], set(), True
for ti, wid in enumerate(enc.word_ids()):
if wid is None or wid in seen:
continue
seen.add(wid)
lab = labels[pred[ti]] # e.g. "O", ",|U", ".", "?|U"
punct = "" if lab[0] == "O" else lab[0]
w = words[wid]
if cap_next or "|U" in lab:
w = w[0].upper() + w[1:]
out.append(w + punct)
cap_next = punct in (".", "?", "!")
return " ".join(out)
print(punctuate("hej mit navn er emil hvordan går det i dag"))
For texts longer than 256 tokens, use the sliding-window helper pnc_infer.py in this repository
(128-token windows with a 12-word overlap, every input word kept in order):
python pnc_infer.py syvai/danish-pnc "hej mit navn er emil hvordan går det i dag"
With ONNX Runtime
pnc.onnx (FP32) and pnc.int8.onnx (dynamic int8, about 22 MB) take input_ids, attention_mask and
token_type_ids and return logits of shape [batch, seq, 10].
import onnxruntime as ort, numpy as np
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("syvai/danish-pnc")
sess = ort.InferenceSession("pnc.int8.onnx")
enc = tok("hej mit navn er emil".split(), is_split_into_words=True, return_tensors="np")
logits = sess.run(None, {k: enc[k] for k in ["input_ids", "attention_mask", "token_type_ids"]})[0]
pred = logits[0].argmax(-1) # map with config.json id2label, first sub-token of each word
Labels
Ten labels, one per word (predicted on the word's first sub-token): punctuation after the word
(O none, ,, ., ?, !) combined with |U when the word should start with an uppercase letter.
The model produces an initial capital only; it cannot restore mixed case such as iPhone.
Training data
All data was streamed from the Hugging Face Hub; nothing was used in full. Sources were chosen for modern, human-punctuated Danish and mixed with the following sampling weights:
| Source | Weight | Register |
|---|---|---|
Danish Dynaword opensubtitles |
0.24 | Film and TV dialogue |
fineweb-2 dan_Latn |
0.24 | Web text |
Dynaword ft |
0.12 | Folketinget transcripts |
Dynaword wikipedia |
0.08 | Encyclopaedia |
Dynaword ep |
0.07 | Europarl transcripts |
Dynaword hest |
0.07 | Forum posts |
Dynaword nordjyllandnews, tv2r |
0.09 | News |
Dynaword mosel_voxpopuli |
0.05 | Parliament speech transcripts |
Dynaword danske-taler, wiki-comments |
0.04 | Speeches, talk pages |
Text was lowercased and stripped of punctuation to form the input; the removed punctuation and original
capitalisation form the labels. Speaker tags, dialogue dashes, quotes and URLs were removed; semicolons map to
commas; colons map to a full stop or comma depending on the following word; abbreviations and ordinals (kr.,
8. oktober) are not treated as sentence ends. Documents with implausible punctuation density, heavy
capitalisation or pre-1948 orthography were dropped. Training windows were 16 to 180 words starting at random
offsets so the model learns to handle chunks that do not begin at a sentence start.
Training: ElectraForTokenClassification from jonfd/electra-small-nordic, 150k steps, batch 32, max length
256, AdamW lr 1e-4 with 3k warmup and cosine decay, weight decay 0.01, fp32 on an Apple M4. The data pipeline
and training code are in data.py in this repository.
Evaluation
Held-out set of 1,750 windows (24 to 160 words) from documents excluded from training by id hash, spanning all
eleven sources. The reference model was scored on exactly the same inputs with the same code
(eval_results.json has the raw numbers, eval_log.jsonl the hourly training curve).
| Metric | RyeAI/ekko-pnc | This model |
|---|---|---|
| Sentence boundary F1 | 0.722 | 0.859 |
| Boundary precision | 0.735 | 0.874 |
| Boundary recall | 0.710 | 0.844 |
| Comma F1 | 0.670 | 0.787 |
| Full stop F1 | 0.696 | 0.825 |
| Question F1 | 0.479 | 0.756 |
| Exclamation F1 | 0.000 | 0.271 |
| Casing F1 | 0.797 | 0.898 |
| Token accuracy | 0.875 | 0.927 |
Sentence boundary F1 per source:
| Source | RyeAI/ekko-pnc | This model |
|---|---|---|
| opensubtitles | 0.724 | 0.892 |
| fineweb2 | 0.654 | 0.776 |
| ft | 0.794 | 0.948 |
| ep | 0.834 | 0.956 |
| wikipedia | 0.756 | 0.894 |
| hest | 0.579 | 0.682 |
| nordjyllandnews | 0.782 | 0.885 |
| tv2r | 0.807 | 0.876 |
| mosel_voxpopuli | 0.729 | 0.733 |
| danske-taler | 0.736 | 0.771 |
| wiki-comments | 0.716 | 0.835 |
Caveats. This eval set shares corpora and labelling conventions with the training data, which favours this
model. RyeAI/ekko-pnc was tuned for continuous meeting speech and evaluated on human-reviewed meeting transcripts,
which were not available here. The mosel_voxpopuli row, the only genuine speech-transcript source, is the
fairest single comparison. Exclamation marks are rare in the eval set and that score is noisy.
Limitations
- Punctuation is predicted from words alone; the model cannot hear pauses or intonation.
- Question marks and exclamation marks are the weakest classes.
- Capitalisation is initial-uppercase only. Keep existing capitals from the input if you have them.
- Trained on edited transcripts and written text; spontaneous conversational speech with disfluencies is under-represented.
License
CC BY 4.0 (see LICENSE). The base model jonfd/electra-small-nordic and the training sources
(Danish Dynaword, CC0; fineweb-2, ODC-By) carry their own terms.
- Downloads last month
- 22
Model tree for syvai/danish-pnc
Base model
jonfd/electra-small-nordic