This collection gather all the metrics used for the evaluation of the datasets in GeoBenchLLM.
Rodrigo Ferreira Rodrigues
rfr2003
AI & ML interests
None yet
Organizations
GeoBenchLLM Metrics
This collection gather all the metrics used for the evaluation of the datasets in GeoBenchLLM.
- RunningAgents
MCQ_eval
🚀Evaluate multiple-choice question predictions
- SleepingAgents
Coord_eval
🚀Evaluate coordinate predictions with a simple Gradio UI
- SleepingAgents
NY_POI_evaluate
🚀Evaluate your model on New York POI dataset
- SleepingAgents
Keywords_evaluate
🚀Evaluate keyword extraction with precision, recall, F1 scores
GeoBenchLLM
The collection for the GeoBenchLLM including the benchmark itself and the metrics used.
spaces 7
Sleeping
Agents
Path_Planning_evaluate
🚀
Evaluate path planning results with performance metrics
Sleeping
Agents
regression_evaluate
🚀
Evaluate regression model predictions with key metrics
Sleeping
Agents
Place_gen_evaluate
🚀
Evaluate your generated place images
Running
Agents
MCQ_eval
🚀
Evaluate multiple-choice question predictions
Sleeping
Agents
NY_POI_evaluate
🚀
Evaluate your model on New York POI dataset
Sleeping
Agents
Keywords_evaluate
🚀
Evaluate keyword extraction with precision, recall, F1 scores
models 0
None public yet