Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
|
bisonbear's submissions
login
1.
I compared Opus 4.8 vs. Opus 5 on 25 of my tasks to see what the difference was
(
stet.sh
)
19 points
by
bisonbear
10 days ago
|
past
|
discuss
2.
I compared 5 popular token saving methods in Codex and found that none delivered
(
stet.sh
)
2 points
by
bisonbear
37 days ago
|
past
3.
I ran Sonnet 5 vs. Opus 4.8 head to head on 24 tasks to see what's different
(
stet.sh
)
1 point
by
bisonbear
51 days ago
|
past
4.
I evaluated GLM 5.2 against the frontier on tasks from real repos
(
stet.sh
)
2 points
by
bisonbear
76 days ago
|
past
|
2 comments
5.
I benchmarked Opus 4.8 vs. GPT 5.5 on 2 open source repos
(
stet.sh
)
3 points
by
bisonbear
3 months ago
|
past
6.
I used autoresearch to improve my AGENTS.md, measured against real tasks
(
stet.sh
)
8 points
by
bisonbear
3 months ago
|
past
|
7 comments
7.
A brief investigation into the GPT-5.5 regression claims
(
stet.sh
)
1 point
by
bisonbear
3 months ago
|
past
8.
The Opus 4.7 reasoning curve - Medium is the best default?
(
stet.sh
)
1 point
by
bisonbear
3 months ago
|
past
9.
GPT-5.5 low vs. medium vs. high vs. xhigh: the reasoning curve on 26 real tasks
(
stet.sh
)
2 points
by
bisonbear
3 months ago
|
past
10.
GPT-5.5 vs. GPT-5.4 vs. Opus 4.7 on 56 real coding tasks from 2 open source repo
(
stet.sh
)
4 points
by
bisonbear
4 months ago
|
past
11.
I ran Opus 4.7 vs. Old Opus 4.6 vs. New Opus 4.6 on 28 Zod tasks
(
stet.sh
)
2 points
by
bisonbear
4 months ago
|
past
12.
Coding evals are broken. CI is green while AI code quality goes unmeasured
(
stet.sh
)
1 point
by
bisonbear
4 months ago
|
past
13.
Agents.md is the highest-leverage code you're not testing
(
stet.sh
)
1 point
by
bisonbear
4 months ago
|
past
14.
Your AI coding benchmark is hiding a 2x quality gap
(
stet.sh
)
3 points
by
bisonbear
5 months ago
|
past
15.
Things I Learned at the Claude Code NYC Meetup
(
benr.build
)
2 points
by
bisonbear
7 months ago
|
past
16.
Claude vs. Codex in the Messy Middle
(
benr.build
)
1 point
by
bisonbear
8 months ago
|
past
17.
Spacetime as a Neural Network
(
benr.build
)
11 points
by
bisonbear
8 months ago
|
past
|
5 comments
18.
One agent isn't enough
(
benr.build
)
18 points
by
bisonbear
8 months ago
|
past
|
2 comments
19.
Context Engineering: The New Skill for Working with AI Agents
(
benr.build
)
1 point
by
bisonbear
10 months ago
|
past
20.
The New Math of Building with AI
(
benr.build
)
2 points
by
bisonbear
10 months ago
|
past
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: