πŸ€– AI Agent Hackathon Build With AI Agents

Build, experiment, and ship working AI-agent projects with developers from across Philadelphia.

πŸ“…
Sept. 20 & 22
πŸ“
Philadelphia
πŸ‘₯
Teams or Solo
πŸ’»
In Person
πŸš€ Register
Register on Luma β†’
πŸ‘₯ Meetup
RSVP on Meetup β†’
πŸ“ Hackathon Overview

Join developers from across Philadelphia for a hands-on hackathon focused on building with AI agents.

Expect practical challenges, working prototypes, and experimentation with agentic systems. You can arrive with a team, meet collaborators at the event, build independently, or join without a finished project idea.

πŸ‘₯ Come with a team 🀝 Find collaborators
πŸ’» Build independently πŸ’‘ No finished idea required
πŸ‘€ Who Should Attend?

Developers, engineers, technical builders, and anyone interested in experimenting with AI-agent systems.

Whether you're experienced with agent frameworks or just starting to explore them, the goal is to build, test ideas, learn from others, and leave with something working.

🧠 What You Can Build With
Coding Agents Multi-Agent Systems Tool Use
Automation Agent Memory & Context Evaluation & Reliability
AI Safety Open-Source Frameworks Real-World Applications
πŸŽ’ What to Bring
πŸ’» Laptop & charger
πŸ› οΈ Your development setup
πŸ”‘ API keys / tools you plan to use
πŸ“… Event Schedule

Two days. Two locations. One hackathon.

DAY 1 β€” BUILD DAY
September 20 β€’ Pennovation Center
3401 Grays Ferry Avenue, Philadelphia, PA 19146
TIME EVENT
9:00 AM – 9:40 AM Registration & Check-In
10:00 AM – 10:15 AM Opening Ceremony
10:15 AM – 10:30 AM Challenge Track Introductions
10:32 AM – 10:35 AM Quirq
10:37 AM – 10:40 AM The Code Registry
10:42 AM – 10:48 AM GalaxyGate & Runpod
10:50 AM – 10:53 AM Human Standard
10:55 AM - 10:58 PM Naftiko
11:00 AM – 11:08 AM Open Track / Bring Your Own Project
11:10 AM – 6:00 PM πŸš€ Hackathon Build Time
DAY 2 β€” AWARDS & COMMUNITY NIGHT
September 22 β€’ Cesium
601 Walnut Street, Suite 250S, Philadelphia, PA 19106
TIME EVENT
6:00 PM Arrival & Check-In
6:15 PM – 7:00 PM 🎀 Featured Speakers & Lightning Talks
7:00 PM – 7:30 PM πŸ† AI Agent Hackathon Awards
7:30 PM+ πŸ• Food & Networking
πŸ† Challenges, Prizes & Judging

Challenge tracks and awards are being finalized with participating partners.

🚧 Final challenge details are coming soon.

Participating sponsors are finalizing their challenge prompts, judging criteria, and prize details. This page will be updated as each track is confirmed.
πŸš€ Ready to Build?

Join developers from across Philadelphia and build something with AI agents.

Register on Luma β†’    |    RSVP on Meetup β†’

The Quirq Challenge: build an agentβ€”or try to break its workspace.

Agents can run for hours, consume tokens, modify files, and call toolsβ€”but it is often difficult to see what they actually did or determine whether they stayed within their intended boundaries. That is the problem Quirq works on.

This challenge has two tracks: Build It and Break It.

Quirq: Build It

Build any useful agentic system and run it through Quirq, using either the managed cloud at app.xo.builders or the local runtime. You may use Claude Code, Codex, OpenClaw, Hermes, n8n, your own agent, or another runtime.

Your project can solve any problem and may also compete in another hackathon track. What matters here is that your agent performs meaningful work and that its activity is visible through the xo-space: its environment, sessions, actions, files changed, costs, and results.

You are not only being judged on whether your agent works. You are also being judged on whether people can understand what it did, evaluate its output, and trust how it operated.

Judging criteria:

  • Agentic Design & Execution β€” 30%
  • Observability & Space Display β€” 25%
  • Impact & Usefulness β€” 25%
  • Reliability & Safety β€” 20%

Prize: $500 cash plus an equivalent amount in Quirq credits. One winner.

Quirq: Break It

Test the security and integrity of your own cloud-managed Quirq workspace at app.xo.builders.

Look for ways to escape the intended workspace scope, conceal agent activity from the xo-space, misreport usage, bypass controls, or otherwise cause the workspace to behave in an unsafe or unexpected way.

Each finding must be submitted privately and include clear reproduction steps, an explanation of its impact, and a suggested fix. Do not test other participants’ workspaces, disrupt shared infrastructure, or target Code & Coffee systems, the local Quirq installation, or third-party model providers.

Validated findings earn points based on:

  • Severity and impact
  • Reproducibility
  • Novelty
  • Scope and systemic impact
  • Quality of the report

Bounty: Up to $500 total may be distributed across accepted findings. Awards depend on the quality and severity of validated findings; the full bounty is not guaranteed. Invalid, duplicate, low-impact, or non-reproducible reports may receive reduced points or no payout.

THE CODE REGISTRY CHALLENGE: BUILD FAST, THEN PROVE IT HOLDS UP.

Today you will generate a lot of code, and most of it will be written by an agent. Almost none of it will be read by a human. That is the problem we work on.

The Code Registry analyses a codebase and returns a Code Score out of 1,000 across three pillars: Security, Dependencies and Quality. There is no LLM anywhere in the analysis pipeline. It is deterministic, so the same code returns the same score every time.

This track runs across the whole event. Build whatever you want, in any language, for any other track. Then run your project through The Code Registry before the cutoff. The highest final Code Score wins.

You are not being judged on whether your agent works. You are being judged on what your agent produced.

Prize: a 12-month Code Registry plan covering up to 400,000 lines of code, listed value $3,600. One winner.

HUMANSTANDARD TRACK: BUILD FOR HUMAN MUSIC

Description:

HumanStandard is a forensic AI music detection company. Our API takes a track and returns a signed verdict on whether the audio was generated by AI, performed by humans, or is a hybrid of the two, along with the evidence behind that call. Distributors, DSPs, and rights holders use it to comply with the EU AI Act, DDEX labeling standards, and platform policies on AI content.

 

For this track, build anything that uses the HumanStandard detection API to solve a real problem in music. Some directions to get you started: a submission screener for a label or playlist curator, a browser extension that flags AI tracks on streaming pages, an artist tool that generates proof of human authorship, a catalog auditing dashboard, a Discord bot for music communities, or an agent that investigates suspicious releases end to end. We care most about ideas that a real person in the music industry would actually want to use.

 

This track is built for people who want to go into music tech. Beyond credits and a fellowship, the winning team gets direct introductions from Rasha Rahman, HumanStandard's founder, to companies across the music industry, and a working session on turning the project into something real. If you are looking for a way into this world, this is the track for you.

 

What you get during the hackathon:

  • A HumanStandard API key issued to your team on the morning of the hackathon, preloaded with 200 credits. One credit is one song scan. Keys are generated from the team email you provide at registration, so make sure that email is one your team actually checks.

  • API documentation and a quickstart you can hand to a coding agent.

  • Arian from HumanStandard will be on site all day to help with the API and answer questions. Rasha Rahman, HumanStandard's founder, will also be available throughout the hackathon and will hold an office hour window where you can bring your project for a critique.

Requirements

QUIRQ SUBMISSION REQUIREMENTS

All Quirq submissions must include:

  • A working prototype

  • A clear project description

  • A link to the project repository

  • A demo video of no more than three minutes

  • A list of the tools, models, frameworks, and runtimes used

  • The names of all team members

  • Setup and testing instructions

Quirq: Build It

Prize

$500 total prize pool, with up to $50 awarded per winning agent.

Each hacker may submit multiple agents. Awards are evaluated at the agent level, so one hacker may receive multiple awards. A hacker with 10 winning agents may win the entire $500 prize pool.

Submission Requirements

Your submission must:

  • Run on Quirq using either the managed cloud at app.xo.builders or a local installation.

  • Show the xo-space during the demo video.

  • Clearly display the agent’s environment, activity, sessions, actions, files changed, costs, and results.

  • Explain the problem being solved and the work completed by each submitted agent.

  • Identify any safeguards, permission controls, evaluations, or failure-handling mechanisms used.

  • Link any relevant xo-space pull request and GitHub Discussion thread when submitting an open-source contribution.

Other agent runtimes are allowed. However, displaying the system’s activity through the xo-space accounts for 25% of the judging score.

Each agent should be submitted and demonstrated clearly enough to be evaluated independently.

Quirq: Break It

Submit one private report for each finding. Each report must include:

  • A clear title and summary

  • The affected feature or component

  • Step-by-step reproduction instructions

  • Evidence such as screenshots, logs, or a short video

  • The actual and expected behavior

  • An explanation of the security or operational impact

  • A suggested fix or mitigation

  • Any relevant environmental or configuration details

All testing must be limited to your own cloud-managed workspace on app.xo.builders. Reports must be submitted privately through the designated Quirq Discord channel before the submission cutoff.

Do not test other participants’ workspaces, disrupt shared infrastructure, publicly disclose findings before review, or target Code & Coffee systems, local Quirq installations, or third-party model providers.

THE CODE REGISTRY
Submission and eligibility requirements

  • Standard Devpost submission by the event deadline
  • One entry per team
  • Repository must be public. Private repositories cannot be cloned for analysis and will not be scored
  • Stop building at 17:30. The final 30 minutes are for registering with The Code Registry, creating a project, and starting the repository sync
  • Hard deadline: the sync must be started before 18:00. Anything started after that is not considered
  • Analysis does not need to complete by 18:00. It runs after you leave, and results are collected before the awards on 22 September
  • Teams must log their team name, repository URL and Code Registry account email with us at the table before leaving
  • Participants will not see their own Code Score during the event. Scoring is carried out by The Code Registry after the build window closes
  • Participants must be over the age of majority in their country of residence (per main event rules)
  • Submitting a project constitutes consent for The Code Registry to analyse the submitted code for judging purposes
  • TCR staff, contractors and their immediate families are not eligible

THE HUMAN STANDARD SUBMISSION REQUIREMENTS

Start of hacking period: 09/20/2026, 10:00 AM (already set)

Deadline: 09/20/2026, 6:00 PM (already set)

Final Reminders (shown right before submit):

Before you submit to the HumanStandard track, make sure you have included the following. Submissions missing any of these will not be judged for the track.

  1. A public code repository link.

  2. A demo video of three minutes or less showing the project working end to end.

  3. A short written description of the problem you are solving and who in the music industry would use it.

  4. An explanation of how your project calls the HumanStandard API and what it does with the results.

  5. Evidence of at least one real API call (a screenshot of a response or a log).

  6. Every team member listed on the submission.

You must opt in to the HumanStandard prize on the submission form to be considered.

File Upload: Off (repo link and video are enough). Require videos: On.

 

Hackathon Sponsors

Prizes

$1,400+ in prizes
+ other prizes
GalaxyGate
1 winner

GalaxyGate 1000$ value compute credits

Quirq Bounty - Break it!
$500 in cash
1 winner

Quirq - Build it!
$500 in cash
1 winner

Prize: $500 cash + equivalent credits β€” guaranteed payout

THE CODE REGISTRY CHALLENGE
1 winner

$3600 value prize, year of up to 400k lines of code. Full visibility and deterministic reporting on security, tech dept, compliance, AI-ROI, third party dependencies, and licensing within your code repository.

Hackathon idea is writing code and whoever gets the highest code score in that time wins the prize!

HumanStandard Track Winner
1 winner

Prize breakdown:

The winning team receives the HumanStandard Fellowship, a package of four things:

10,000 API credits and expanded model access. Keep building after the hackathon with 10,000 credits plus access to detection models beyond the base API, including the hybrid model (AI vocals over human instrumentals) and the provenance model.

A feature to 50,000 people. Rasha Rahman, HumanStandard's founder, will present and share your project with HumanStandard's community of over 50,000 followers on Instagram and TikTok, an audience made up of musicians, producers, and people working in the music industry.

60 minutes of bookable time with Rasha. Use it however you want: career advice, startup critique, marketing and go to market advice, or help turning the project into something real. Email rasha@hsverify.com whenever you are ready to schedule, in one session or several.

Industry connections. After the session, Rasha will make personal introductions to a few people at the largest music and music tech companies in the world, chosen based on what you built and where you want to go.

HumanStandard Track Runner Up
1 winner

5,000 HumanStandard API credits to keep building after the hackathon.

Bring your own Agent Project
$400 in cash
1 winner

Devpost Achievements

Submitting to this hackathon could earn you:

Judges

Tony Siu

Tony Siu
Agent Software Engineer ||| - Zoominfo

satya veerendra Vegulla

satya veerendra Vegulla
Senior Engineering Manager/Principal Architect

Chase Lauer

Chase Lauer
Senior Network Engineer - GalaxyGate

Suraj Sharma

Suraj Sharma
Founder - Quirq

Satvik Bhasin

Satvik Bhasin
Platform Engineer - Geico

Aditi Patodiya

Aditi Patodiya
Senior Software Engineer - Amazon

Chandler Samuels

Chandler Samuels
Associate Director of Solutions Engineering - Curotec

Arun Aalla

Arun Aalla
Manager - Deloitte Digital

Bharat Patel

Bharat Patel
Lead Software Engineer, Data & AI at Intuit

Richie Singh

Richie Singh
Venture Capital & AI @ Mucker Capital | AI Researcher & Engineer

Muhammad Rashid
Founder & CEO of CureOn

DakShith Ragupathi

DakShith Ragupathi
Chief Business Development Officer, CureOn

Hasan Shameer

Hasan Shameer
Software Engineer - Bristol Myers Squibb

Judging Criteria

  • Bring Your Own Project - Technical Execution
    How well does the project work? Is the implementation technically sound and functional?
  • Bring Your Own Project - Agentic Design
    How meaningfully does the project use agents, tool calling, memory, planning, multi-agent workflows, or autonomous execution?
  • Bring Your Own Project - Innovation & Creativity
    How original is the idea, approach, or use of agent technology?
  • Bring Your Own Project - Impact & Usefulness
    Does the project solve a meaningful problem or provide clear value to users?
  • Bring Your Own Project - Reliability & Safety
    Does the team account for failure modes, evaluation, permissions, guardrails, or other reliability and safety considerations?
  • Bring Your Own Project - Demo & Completeness
    How clearly does the team demonstrate what they built, and how complete is the working prototype?
  • Human Standard - Does it work?
    The project runs and does what the team claims. The HumanStandard API is actually integrated and the verdict data is used in a meaningful way, not just displayed.
  • Human Standard - Is it a real problem?
    The team identified a genuine pain point for someone in the music industry, whether an artist, label, distributor, curator, platform, or fan. Bonus points for teams that can name who would use this and why they would care.
  • Human Standard - Product quality
    The user experience is thoughtful. Verdicts are communicated honestly and with appropriate uncertainty rather than as a binary "this is AI" claim. The project feels like the start of something a person could ship, not just a proof that the API responds.
  • Human Standard - Music tech ambition
    The team shows real curiosity about the music industry and a desire to keep building in it. This matters because the prize is largely career oriented: introductions, mentorship, and a fellowship are most valuable to people who want to pursue this path.
  • Quirq Build it - 30% weight; Agentic Design & Execution
    Meaningful use of agents, including tool calls, memory, planning, multi-agent collaboration, and autonomy. Most importantly: does the system actually work?
  • Quirq Build it - 25% Observability & Space Display
    How clearly the demo shows the environment, agents, and their activityβ€”including files touched, sessions, and costs. The Quirq space must be part of the demo.
  • Quirq Build it - 25% Impact & Usefulness
    Whether the project solves a real problem for a real user.
  • Quirq Build it - 20% Reliability & Safety
    How well the project handles failure modes, scoped permissions, guardrails, and evaluation.
  • Quirq Break It Bounty
    Quirq Red Team Bounty is evaluated separately using a points-based system: Severity and impact, Reproducibility, Novelty, Scope or systemic impact, Quality of the report
  • The Code Registry Bounty
    Eligible projects will be analyzed using The Code Registry. The participant whose hackathon code receives the highest final Code Score will win.

Questions? Email the hackathon manager

Tell your friends

Hackathon sponsors

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.