LIFEHUBBER
Choose theme

AI Resources

General365

GitHub stars: 87 GitHub forks: 6 Declared license: MIT: MIT Last pushed April 14, 2026: Pushed 5mo ago
Stats from GitHub

General365 is a reasoning benchmark from meituan-longcat for comparing how language models handle varied problems with background knowledge designed to stay within K-12 scope.

It provides public questions and grading code, so you can collect responses from the models you want to compare and examine their results question by question. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.

What it is

Problems, responses and a grader

The project supplies manually curated reasoning problems and an evaluation workflow. It is a benchmark to run with a model, not a model checkpoint or chat application.

Why it stands out

Reasoning without specialist prerequisites

Its design limits the required background knowledge while varying the reasoning challenge. This targets a different question from tests that rely on university-level subject knowledge.

Availability

A public subset, code and paper

The repository links the General365_Public dataset, grading scripts, paper and project page. The authors retain part of the questions as a held-out set; downloading the public data does not give you the full evaluation set.

Why it matters

What makes it useful

The problems test following constraints and logical branches rather than only recalling a specialist fact. Use their individual questions to identify the reasoning task you want to compare; a benchmark score does not establish how a model will handle your own application.

Notable points

What stands out

Grading depends on answer type. Text and choice answers go to a configured model grader; numeric and other symbolic answers use the mathematical verification path. The model-grading request includes the question, reference answer and submitted response. Check that endpoint’s cost and data handling before running it; the whole grading workflow is not offline. Before comparing two average scores, compare the question IDs in their per_question_accuracy results with the IDs you intended to evaluate. The script averages the unique IDs supplied in the response file, and a repeated ID replaces the previous result for that question. A mean alone therefore does not show that every intended question was covered. Align the question sets before treating the two means as a comparison on the same test.

Before using

What to review

The paper describes 365 seed problems and 1,095 variant problems across eight categories. These counts describe the benchmark design, not the number of questions you can download from its public subset.

Variants alter wording or constraints while retaining the target reasoning skill. When interpreting a failure, check which constraints that particular question asks the model to follow.

Held-out questions are the authors’ method for tracking potential contamination, not proof that a model has never encountered the public questions.

Reader fit

Who may find it relevant

Developers comparing saved responses from different models on the same reasoning questions.

Researchers who want to inspect problems, variants and grading choices rather than only a leaderboard position.

Someone evaluating a specific product workflow still needs examples from that workflow; this benchmark does not replace them.

Editorial note

Why LifeHubber lists it

An overall reasoning score can hide which kinds of problems a model struggles with. General365’s project materials tag seed problems by challenge and show results across eight challenge categories, giving readers a way to examine that pattern alongside the average. That is useful when choosing which reasoning failures to investigate next. The publisher’s category analysis is separate from the supplied grading script’s per-question results; do not assume the public download automatically reproduces that analysis.

Source links

Source materials

Reader note

Before relying on this entry

LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.

Advertisements

Advertisements

For project maintainers

Listed here? You can use the badge.

If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.

See what’s moving