In other words: Investigate why and in which ways they fail on unknown knowledge graphs, the most prominent factors leading to failure
Possible directions:
- Detailed error analysis over common causes of failure
- Can failures related to unfamiliartiy with the KG be solved via explicit prior exploration?
- Are failures repeatable, e.g., if asking the same question multiple times, will it always fail in the same way?
- Produce a garbled version of a well-known KG from pre-training (e.g., Wikidata) and investigate changes in agent behavior