QUIRQ JUDGING CRITERIA

Quirq: Build It

Agentic Design & Execution β€” 30%

  • Meaningful use of agents, tools, memory, planning, autonomy, or multi-agent collaboration
  • Quality of the system’s architecture and execution
  • Whether the agent successfully completes its intended work

Observability & Space Display β€” 25%

  • Clear visibility into the agent’s environment and activity through the xo-space
  • Ability to inspect sessions, actions, files changed, costs, handoffs, and results
  • How effectively the project helps users understand what the agent did

Impact & Usefulness β€” 25%

  • Importance of the problem being addressed
  • Practical value for a real user or organization
  • Quality and usefulness of the delivered result

Reliability & Safety β€” 20%

  • Handling of errors and failure modes
  • Appropriate permissions, safeguards, and guardrails
  • Testing, evaluation, and consistency of results
Quirq: Break It

Break It submissions receive points for each validated finding based on:

Severity and Impact

  • The potential harm or operational consequence of the issue
  • Whether the finding affects confidentiality, integrity, availability, isolation, or reporting accuracy

Reproducibility

  • Whether the issue can be consistently reproduced
  • Quality and completeness of the reproduction steps and supporting evidence

Novelty

  • Whether the finding identifies a new issue
  • Duplicate or previously known findings may receive reduced points or no points

Scope and Systemic Impact

  • Whether the issue affects one feature or represents a broader platform weakness
  • Potential impact across multiple workspaces, users, or workflows

Report Quality

  • Clarity of the explanation
  • Strength of the supporting evidence
  • Accuracy of the impact assessment
  • Usefulness of the proposed fix or mitigation

The highest-scoring participant may be recognized as the Break It winner, but monetary awards depend on the strength of the validated findings. Up to $500 may be distributed across accepted submissions, and the complete bounty pool is not guaranteed.

HUMAN STANDARD JUDGING CRITERIA

Does it work?

The project runs and does what the team claims. The HumanStandard API is actually integrated and the verdict data is used in a meaningful way, not just displayed. A partially working project with a clear core loop beats a broad project where nothing quite finishes.

 

Is it a real problem?

The team identified a genuine pain point for someone in the music industry, whether an artist, label, distributor, curator, platform, or fan. Bonus points for teams that can name who would use this and why they would care.

 

Product quality

The user experience is thoughtful. Verdicts are communicated honestly and with appropriate uncertainty rather than as a binary "this is AI" claim. The project feels like the start of something a person could ship, not just a proof that the API responds.

 

Music tech ambition

The team shows real curiosity about the music industry and a desire to keep building in it. This matters because the prize is largely career oriented: introductions, mentorship, and a fellowship are most valuable to people who want to pursue this path.

 

Tiebreaker guidance for judges: When multiple projects meet all four criteria, favor the one whose team is most likely to keep building after the hackathon and whose idea a music company would most want to see.

THE CODE REGISTRY JUDGING CRITERIA

Primary metric: highest final Code Score (out of 1,000) from the last completed analysis submitted before the cutoff.

Eligibility floor. To be scored, a project must:

  • Be a valid hackathon submission under the main Devpost rules (working prototype, repository, description, demo video)
  • Live in a new public GitHub repository created on 20 September, with hackathon in the repository name
  • Be a working application with a stated goal, not a script or a fragment
  • Include a README covering what the app does, the problem it solves, how to run it, how AI agents are used, and who is on the team
  • Contain at least 1,500 lines of source code across at least 10 source files
  • Declare at least three third-party dependencies in a package manifest
  • Be built during the event window on 20 September, with the Code Registry sync started before 18:00 to be considered

Projects below the floor are not scored. This exists so the prize goes to a real build rather than the smallest possible repository.

Commit timestamps are the check on "built during the event". A repository created that morning with commits running through the day evidences itself.

Tiebreakers, in order:

  1. Highest Security pillar score
  2. Highest Quality pillar score
  3. Judges' view of the working application against its stated goal