Choose theme
AI Resources
General365
General365 is a reasoning benchmark from meituan-longcat for comparing how language models handle varied problems with background knowledge designed to stay within K-12 scope.
It provides public questions and grading code, so you can collect responses from the models you want to compare and examine their results question by question. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.
What it is
Problems, responses and a grader
The project supplies manually curated reasoning problems and an evaluation workflow. It is a benchmark to run with a model, not a model checkpoint or chat application.
Why it stands out
Reasoning without specialist prerequisites
Its design limits the required background knowledge while varying the reasoning challenge. This targets a different question from tests that rely on university-level subject knowledge.
Availability
A public subset, code and paper
The repository links the General365_Public dataset, grading scripts, paper and project page. The authors retain part of the questions as a held-out set; downloading the public data does not give you the full evaluation set.
Why it matters
What makes it useful
The problems test following constraints and logical branches rather than only recalling a specialist fact. Use their individual questions to identify the reasoning task you want to compare; a benchmark score does not establish how a model will handle your own application.
What to know
Where it fits
Collect model answers into a JSONL file with question_id and model_response in each row, place it in model_responses, then run grading.py with that filename. The script writes average_accuracy and per_question_accuracy under grading_results. This separates generating the answers from grading them, so you can retain the responses alongside the results.
Notable points
What stands out
Grading depends on answer type. Text and choice answers go to a configured model grader; numeric and other symbolic answers use the mathematical verification path. The model-grading request includes the question, reference answer and submitted response. Check that endpoint’s cost and data handling before running it; the whole grading workflow is not offline. Before comparing two average scores, compare the question IDs in their per_question_accuracy results with the IDs you intended to evaluate. The script averages the unique IDs supplied in the response file, and a repeated ID replaces the previous result for that question. A mean alone therefore does not show that every intended question was covered. Align the question sets before treating the two means as a comparison on the same test.
Before using
What to review
The paper describes 365 seed problems and 1,095 variant problems across eight categories. These counts describe the benchmark design, not the number of questions you can download from its public subset.
Variants alter wording or constraints while retaining the target reasoning skill. When interpreting a failure, check which constraints that particular question asks the model to follow.
Held-out questions are the authors’ method for tracking potential contamination, not proof that a model has never encountered the public questions.
Reader fit
Who may find it relevant
Developers comparing saved responses from different models on the same reasoning questions.
Researchers who want to inspect problems, variants and grading choices rather than only a leaderboard position.
Someone evaluating a specific product workflow still needs examples from that workflow; this benchmark does not replace them.
Editorial note
Why LifeHubber lists it
An overall reasoning score can hide which kinds of problems a model struggles with. General365’s project materials tag seed problems by challenge and show results across eight challenge categories, giving readers a way to examine that pattern alongside the average. That is useful when choosing which reasoning failures to investigate next. The publisher’s category analysis is separate from the supplied grading script’s per-question results; do not assume the public download automatically reproduces that analysis.
Source links
Source materials
Reader note
Before relying on this entry
LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.
More in Datasets
Keep browsing this category
Explore more datasets.
ParseBench
run-llama/ParseBench
A document parsing benchmark for AI-agent workflows, focused on whether parsed PDFs preserve structure and meaning for downstream evaluation.
UltraData-SFT-2605
openbmb/UltraData-SFT-2605
An OpenBMB supervised fine-tuning dataset with 15,036,178 thinking and non-thinking samples across math, code, knowledge, Chinese, instruction-following, and multilingual configurations, used in MiniCPM5-1B-SFT post-training.
Monitorability Evals
openai/monitorability-evals
An OpenAI evaluation-data release for studying monitorability, with public eval splits, prompt templates, dataset mappings, and metric code from the Monitoring Monitorability paper.
For project maintainers
Listed here? You can use the badge.
If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.