RSA-896 is a RSA challenge number I factored with Claude on September 19, 2026. More details will follow. Briefly, Claude was used to port CADO-NFS to run on GPUs. It orchestrated running on a fleet of up to 2048 GPUs as a low-priority job during unused idle time between regular jobs. The computation ran over a 10-day period and performed about 30 GPU-years of compute time at Anthropic.
This work did not meaninfgully improve the runtime of the General Number Field Sieve (GNFS) algorithm. It does not impact the security of deployed RSA-2048 keys. However, it does demonstrates that RSA-1024 keys are vulnerable to many actors with data center-level fleets of GPUs.
RSA-896 = 4120234369866595438555313653325759481798116998443279828454556264 3387644556524842619809887042316184187926142024718886949256093177 6375033421130982397485150944909106910269861031862704114880866970 5649029036536588674337317208131041051908642547932826013912576240 33946373269391
p = 636606729769440499166579950236036751749912014371509557713570027 508971809534551913252252094954941974952859310861988904737359709 200557919 q = 647218161102195448058768698177623951380616936266986989243011933 572862870905830904361851542450154852431416136790787107595965374 752513489
The polynomial used had degree 6, alpha -11.12, Murphy E 5.293e-10, Res(f,g) = -8N, and was this:
Y0: -13251813291383946856267859335343621268178701 Y1: 73476233331852469840135261 c0: 47116807724228742541747038893184760096436784698070199672 c1: -32945633819741939141253692412171597093977742630162 c2: -2025961670739321318539049635089452233067707 c3: 195344412542670829135241100477177798 c4: 4560157188630514979797057951 c5: -203811845861588319720 c6: -608639391360 skew: 20926672.028
side 0 (rational) side 1 (algebraic, degree 6)
lim 2^32 - 1 2^32 - 1
lpb 37 40
mfb 74 (2 large primes) 120 (3 large primes)
Special-q: prime, on side 1
qmin: 6.0e9
qmax: 4.90e10 (frontier at the stop)
Coverage: about 98.7% of that range was completed. Work in flight
at the stop was dropped.
Sieve area: A = 33, a fixed 2^17 x 2^16 region.
From about 4 hours in, cells with i and j both even were skipped.
That leaves 3/4 of the cells and gave identical relations in testing.
We sieved at lpb 37/40 but built the matrix at 37/39. That dropped
35.0% of the relations.
None. Both sides were sieved over the whole range. Cofactors were factored by ECM on the GPU with B1 = 500 and B2 = 25000. The budget was 35 curves, or 100 for a piece that can only be two large primes. Our batch prefilter stayed off because it does not pay above mfb 96.
Special-q sieved: 1.774e9 (q, root) pairs
Relations found: 37.81 G
Kept as unique: 29.05 G
Per special-q: 21.3 found, 16.4 unique (average)
Those counts cover 267,724 of the 267,969 work blocks. The final
unique set was 29,076,172,342 relations.
Test sieve on the final polynomial, one H100, before the sublattice
change (multiply ms/q by 0.88 for the rate afterwards):
q found/q unique/q ms/q duplicates
6e9 32.1 32.1 380 0.0%
3e10 19.4 13.9 372 28.2%
4e10 17.7 12.2 370 31.0%
5.5e10 16.4 10.7 369 34.5%
Production agreed with this:
part of the run found/q unique/q
first 134 M special-q 29.8 28.3
last 130 M special-q 17.1 11.4
In H100 GPU-seconds, not core-seconds. The host CPU only does the
q-lattice reduction.
definition s per special-q
sieving + cofactorization inside a work unit 0.333
(before the sublattice change) 0.377
median over whole work units, incl. process start 0.340
all allocated GPU time in the sieving window,
divided by special-q completed 0.361
The last figure includes work redone after preemptions. It comes to
177,929 GPU-hours, or 0.022 GPU-s per unique relation.
CPU reference, not like for like: stock las at A = 33 took 290 to 334
core-s per special-q on Graviton3, and about 163 physical-core-s on a
Xeon 8559C. That test used the RSA-260 polynomial with lim 2^31 and
lpb 37/38.
23.2% of relations found (8.76 G of 37.81 G). The siever removed them on the fly. A relation was dropped if a smaller special-q in [qmin, q) would also find it, so the raw set was never stored. The filter's own (a,b) check then found 0 duplicates in the 29.08 G lines.
Input: 18,895,476,871 unique relations at 37/39, plus 7.76 M
free relations, against 20.10 G columns
After purge: 5,913,776,340 x 5,913,776,180, weight 154.6 G
After merge (the matrix solved):
rows 1,482,375,753
columns 1,482,341,556
nonzeros 296,475,150,709 (200.0 per row)
The 32 heaviest columns (20.4 G nonzeros) were held out as a dense
block and handled at the characters step. Counting them, the matrix
has 213.8 nonzeros per row.
Block Wiedemann: m = n = 3584, as 56 sequences of width 64, with
827,234 iterations each.
Total 10.07 days (870,124 s) on up to 256 preemptible nodes of
8 H100 each.
phase wall hours
polynomial selection 14.9
test sieving, launch 2.2
sieving 90.9
filtering 21.9
Krylov 58.8
Krylov check, lingen 40.2
mksol, gather 7.4
characters, square root 5.3
Notes:
- Polynomial selection ran on 128 nodes.
- Sieving includes three whole-fleet outages of 35 to 90 minutes.
- Filtering was about 3.1 h of actual work on 8 nodes. The rest went
to a preemption, a false alarm in a pre-solve audit, and preparing
the matrix for the solve.
- The first half of lingen ran alongside Krylov. The Krylov window
includes a fleet relaunch.
- Lingen was restarted four times.
- mksol lost about 3 h to an out-of-memory failure.