Wanghan Xu commited on
Commit
79e99c1
·
verified ·
1 Parent(s): 7eb5002

Add ResearchClawBench evaluation result

Browse files

Adds ResearchClawBench overall evaluation metadata.

ResearchClawBench is a tool-using, code-execution, file-system-workspace benchmark. The value is the mean score out of 100 over completed ResearchHarness runs, with details available from the ResearchClawBench leaderboard.

.eval_results/researchclawbench.yaml ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ - dataset:
2
+ id: InternScience/ResearchClawBench
3
+ task_id: overall
4
+ value: 18.00
5
+ date: "2026-05-18"
6
+ notes: "ResearchHarness: https://huggingface.co/spaces/InternScience/ResearchHarness; ResearchClawBench: https://huggingface.co/datasets/InternScience/ResearchClawBench; tools enabled; code execution; file-system workspace; completed 38/40 tasks"
7
+ source:
8
+ url: https://huggingface.co/moonshotai/Kimi-K2.6
9
+ name: Model Card
Free AI Image Generator No sign-up. Instant results. Open Now