cohere: support untied head weight - #179
Open
LasseLegarth wants to merge 1 commit into
Open
LasseLegarth wants to merge 1 commit into
LasseLegarth wants to merge 1 commit into
Conversation
Fine-tunes such as syvai/hviske-v5* train the LM head separately from the token embedding. Detect this in convert-cohere.py and write a dedicated head.weight tensor; at runtime, prefer it over the tied embedding when present. Also accept float32 source tensors and rebuild the mel filterbank from librosa when the checkpoint omits the preprocessor buffer.
Contributor
|
Thanks for submitting this. I will take a look soon |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Some Cohere fine-tunes (e.g. syvai/hviske-v5*) train the LM head
separately from the token embedding. The current code assumes they are
always tied, producing wrong logits for these checkpoints.
convert-cohere.py now detects an untied head and writes a dedicated
head.weighttensor. The runtime prefers it over the tied embeddingwhen present.
Scope
AI Assistance
Claude assisted with parts of the implementation. I reviewed and tested
the changes myself.
Validation
Converted and ran syvai/hviske-v5-tiny (untied head) against Danish
speech. Before/after:
I'm running Handy daily with these changes and hviske-v5-tiny as my
dictation model. Works well for Danish input.
Also accepts float32 source tensors and rebuilds the mel filterbank
via librosa when the checkpoint omits the preprocessor buffer.