I've never looked at the LiDAR hardware, but where is the emitter in relation to the receiver. Why would the LiDAR not reflect off of whatever you're blocking it with and return a very short flight meaning it was very close?
I think optics could be used to make each camera see a different image.
A video could show shake, which could be verified against readings from the phone's accelerometer -- but you could just hold it still and claim that it was on a tripod.
Similarly, LiDAR alone will help disqualify cases where someone is just taking a picture of e.g. a landscape target of the Golden Gate, but that it shown on a screen 1 meter away.