Tagged tokenization
1 write-up.
Why a Small Transformer Can't Copy a Word It Hasn't Seen
An 11.9M-parameter transformer scores 83% strict on the domains it trained on and 0/8 on a single unseen noun. Three experiments to find out why, one of which corrected the diagnosis I had already written down, and a zero-parameter mechanism that does the copying 20 times out of 20.