QUIRQ JUDGING CRITERIA
Quirq: Build It
Agentic Design & Execution β 30%
- Meaningful use of agents, tools, memory, planning, autonomy, or multi-agent collaboration
- Quality of the systemβs architecture and execution
- Whether the agent successfully completes its intended work
Observability & Space Display β 25%
- Clear visibility into the agentβs environment and activity through the xo-space
- Ability to inspect sessions, actions, files changed, costs, handoffs, and results
- How effectively the project helps users understand what the agent did
Impact & Usefulness β 25%
- Importance of the problem being addressed
- Practical value for a real user or organization
- Quality and usefulness of the delivered result
Reliability & Safety β 20%
- Handling of errors and failure modes
- Appropriate permissions, safeguards, and guardrails
- Testing, evaluation, and consistency of results
Quirq: Break It
Break It submissions receive points for each validated finding based on:
Severity and Impact
- The potential harm or operational consequence of the issue
- Whether the finding affects confidentiality, integrity, availability, isolation, or reporting accuracy
Reproducibility
- Whether the issue can be consistently reproduced
- Quality and completeness of the reproduction steps and supporting evidence
Novelty
- Whether the finding identifies a new issue
- Duplicate or previously known findings may receive reduced points or no points
Scope and Systemic Impact
- Whether the issue affects one feature or represents a broader platform weakness
- Potential impact across multiple workspaces, users, or workflows
Report Quality
- Clarity of the explanation
- Strength of the supporting evidence
- Accuracy of the impact assessment
- Usefulness of the proposed fix or mitigation
The highest-scoring participant may be recognized as the Break It winner, but monetary awards depend on the strength of the validated findings. Up to $500 may be distributed across accepted submissions, and the complete bounty pool is not guaranteed.
HUMAN STANDARD JUDGING CRITERIA
Does it work?
The project runs and does what the team claims. The HumanStandard API is actually integrated and the verdict data is used in a meaningful way, not just displayed. A partially working project with a clear core loop beats a broad project where nothing quite finishes.
Is it a real problem?
The team identified a genuine pain point for someone in the music industry, whether an artist, label, distributor, curator, platform, or fan. Bonus points for teams that can name who would use this and why they would care.
Product quality
The user experience is thoughtful. Verdicts are communicated honestly and with appropriate uncertainty rather than as a binary "this is AI" claim. The project feels like the start of something a person could ship, not just a proof that the API responds.
Music tech ambition
The team shows real curiosity about the music industry and a desire to keep building in it. This matters because the prize is largely career oriented: introductions, mentorship, and a fellowship are most valuable to people who want to pursue this path.
Tiebreaker guidance for judges: When multiple projects meet all four criteria, favor the one whose team is most likely to keep building after the hackathon and whose idea a music company would most want to see.
THE CODE REGISTRY JUDGING CRITERIA
Primary metric: highest final Code Score (out of 1,000) from the last completed analysis submitted before the cutoff.Eligibility floor. To be scored, a project must:
- Be a valid hackathon submission under the main Devpost rules (working prototype, repository, description, demo video)
- Live in a new public GitHub repository created on 20 September, with hackathon in the repository name
- Be a working application with a stated goal, not a script or a fragment
- Include a README covering what the app does, the problem it solves, how to run it, how AI agents are used, and who is on the team
- Contain at least 1,500 lines of source code across at least 10 source files
- Declare at least three third-party dependencies in a package manifest
- Be built during the event window on 20 September, with the Code Registry sync started before 18:00 to be considered
Projects below the floor are not scored. This exists so the prize goes to a real build rather than the smallest possible repository.
Commit timestamps are the check on "built during the event". A repository created that morning with commits running through the day evidences itself.
Tiebreakers, in order:
- Highest Security pillar score
- Highest Quality pillar score
- Judges' view of the working application against its stated goal
