upvote
I do this with a summary prompt + a JSON config. The config keys are the full name of the thing that gets commonly mistranscribed, and the values include a phonetic spelling of the name and a n "entity id" which is just the name of a markdown file that provides information about the entity, which the LLM can use to infer what should have been referenced if it's not clear.

My prompt is pretty simple:

    - Intro of the purpose (faithful, high-quality summaries of meeting transcripts useful for readers who did not attend the meeting)
    - Some basic rules
        - Summarize only what is actually said in the transcript.
        - If a speaker's name isn't clear from the transcript, describe them generically (e.g. "a participant") rather than guessing
        - If the transcript is a fragment or cuts off, say so rather than inventing a resolution
        - Only state a causal or explanatory link between two facts (e.g. why someone is absent) if that link is clearly stated in the transcript
        - Preserve relative time references precisely rather than flattening them. If a speaker says something will happen "in two hours" don't reword it as "today" or "same day"
        - Maintain the original voice and framing to keep original intent
        - Attempt to infer meeting attendance based on who actually spoke or was spoken to
        - Prefer plain, direct sentences over compound or clever phrasing
        - Respond directly with the summary. Do not preface it with phrases like "Here is a summary" or "Based on the transcript".
    - A preferred structure for the summary, with examples
    - An instruction to attempt correction of mis-transcribed words based on the included JSON configuration, but only if the certainty is high. If unsure, leave as-is, but add (spelling?) after the word to flag it.
    - The JSON config as a string.
I invoke this with a skill that can also update the JSON config if I identify new mistranscriptions, so it gets updated regularly.
reply