Goal
Compare observed GPT and Gemini behavior in our route-navigation task.
Description
Write a project-specific comparison of GPT and Gemini based on actual dashboard/evaluation outputs.
The comparison should focus on practical behavior in our setup, not broad claims about the models in general.
Possible comparison points:
- route validity
- exact shortest-path match
- JSON format reliability
- tendency to hallucinate nodes or edges
- ability to use SSAL graph context
- sensitivity to prompt template changes
- explanation quality
- finish status / truncation behavior
- consistency across similar prompts
- common failure modes
Literature / background review
Include a small literature/background review on comparative LLM evaluation and model-to-model benchmarking.
The goal is to make the GPT/Gemini comparison methodologically careful.
The writeup should avoid broad claims such as “GPT is better at routing” unless supported by our data. Instead, it should frame the comparison as project-specific behavior under our route-navigation setup.
Useful angles:
- comparative model evaluation
- benchmark design and small-sample limitations
- prompt sensitivity across models
- output-format reliability
- model behavior on reasoning or spatial tasks
Expected output:
- 3–5 relevant sources
- short note on how to compare models without overclaiming
- report-ready limitation paragraph
Suggested work
Acceptance criteria
Goal
Compare observed GPT and Gemini behavior in our route-navigation task.
Description
Write a project-specific comparison of GPT and Gemini based on actual dashboard/evaluation outputs.
The comparison should focus on practical behavior in our setup, not broad claims about the models in general.
Possible comparison points:
Literature / background review
Include a small literature/background review on comparative LLM evaluation and model-to-model benchmarking.
The goal is to make the GPT/Gemini comparison methodologically careful.
The writeup should avoid broad claims such as “GPT is better at routing” unless supported by our data. Instead, it should frame the comparison as project-specific behavior under our route-navigation setup.
Useful angles:
Expected output:
Suggested work
Acceptance criteria