So, what you do is you recognize every ticket has a somewhat variable “actual effort”; and, if you’ve been honest in approximate effort pointing, you’ll know your team (or your own) velocity.
From there you can run Monte Carlo simulations - say a few hundred thousand, and get a pretty good estimate of actual time spent.
I’ve seen it work before with shocking accuracy.
estimate(human_estimator, task_description, world_state) -> numeric_effort
Assume that for various practical reasons, we've decided it's one of the best functions out there. How do we use it effectively, especially when it has noise, and drifts over time with unseen changes to the human_estimator and the hideously complex world_state?
A popular option is to run it multiple times with different person/task combinations, putting a projected number on to each task. Afterwards, the tasks finished in sampling period ("sprint") become a quantifiable total for that period ("velocity").
Do the same process again with the next set of tasks, and you can figure out which ones are likely to fit if the velocity doesn't change much. If you know the velocity will change due to losing staff or vacation days... well, we apply a multiplier and hope for the best.
Trying to "fix" the meaning of points is maladaptive, because they reflect many changing things which are outside our control and can't be independently measured.