Language models have a critical handoff point where they transition from using query routing information to relying on internal knowledge—this happens at different layers across models and reveals how they internally organize and access information.
This paper investigates how large language models retrieve and use internal knowledge when answering questions by analyzing how different layers process query information versus stored knowledge.