Research record
Reports and supporting evidence
Read the study records behind our findings, including unsuccessful tests and work that remains incomplete.
Dynamical Synthesis: Learning through Interaction
The first paper brings together the framework, completed experiments and remaining gaps. The evidence bundle includes saved mathematics replies and review decisions.
Study reports
11 September 2026
Research paper and supporting data
The paper draws on our existing experiments to examine where performance improved, where tests failed and which questions remain open. We provide supporting data and analysis files so readers can review the evidence behind these findings.
Research paper
7 September 2026
Preparation for the learning comparison
The latest round collected 175 of 336 planned measurements, while two earlier rounds collected all their data but found too few suitable questions. The main comparison is unfinished, so we cannot tell whether the approach improves learning.
Incomplete study
7 September 2026
Performance after further training
Prediction scores worsened on new and earlier test examples. Across 42 questions tested twice, 12 answers improved, 17 worsened and 55 stayed the same. Seven losses involved earlier skills, and neither update was adopted.
Completed comparison
7 September 2026
Separate prediction and response updates
The plan called for separate prediction and answer updates on a copy of the system, reusing earlier examples. Test examples would stay out of training, and updates would need to pass checks for improvement and preserving earlier skills before adoption.
Study plan
6 September 2026
Selecting study questions
The plan would first look for questions that produced consistent results before training, then divide them into training and test sets. This report describes how questions would be chosen; later reports give the results.
Study preparation
6 September 2026
Rule correction with feedback
All 70 planned requests finished; two replies hit the length limit. Records showed a correct rule and successful reuse in one of six worlds with checked feedback, none with self-checking and two with supplied rules. No wrong initial rule was corrected, and conflicting records prevented a final decision. An earlier unscored request is reported separately.
Incomplete comparison
6 September 2026
Learning and reusing hidden rules
The system identified 2 of 6 hidden rules correctly and gave 3 correct answers across 12 later questions using its saved rules. It did not pass the full learning and reuse test, which kept model weights fixed and supplied the saved rules in the prompt.
Completed comparison
6 September 2026
Reasoning settings with supplied rules
With rules provided for 12 questions, the prediction adapter scored 9 with extra reasoning and 2 without it; the answer adapter scored 7 and 2. Extra reasoning also required more generated text. The study tested use of given rules, without measuring rule discovery.
Diagnostic comparison
5 September 2026
Structured and direct answers
On 8 selected questions, the trained adapter passed 4 with forced output formatting and 4 without it; a control trained with mismatched answers passed none. Direct numerical answers scored 3 and 2 respectively. Structured replies used the checker for arithmetic, so the comparison changed both format and calculation support.
Diagnostic comparison
5 September 2026
Initial tests of rule learning
Across 8 simulated worlds, no initial rule was correct and one became correct after feedback. Using its own records, the system passed none of 16 later test sequences; with the original observations, it passed one. These results did not meet the requirements for starting the planned delayed test.
Incomplete study
5 September 2026
Harder mathematics comparison
Each model answered 27 problems three times. The trained adapter and a control trained with mismatched answers each passed 2 of 81 replies; the earlier prediction adapter passed 3. The study did not show a gain over the control. Difficulty and answer requirements both changed from the earlier study, so the cause of the difference remains unclear.
Completed comparison
4 September 2026
Training on familiar problem types
The trained adapter passed 43 of 64 questions with new numbers in 16 problem types used in training, while the base model and a control trained with mismatched answers passed none. Accepted answers passed exact mathematical checks and a separate AI review, but the finding remains limited to these problem types and conditions.
Completed comparison
Access to further evidence
Some reports reference records in the research archive. The public bundle contains the released mathematics evidence; other records can be requested through Dilate, subject to privacy and review.
Loopseed
A research programme at Dilate Technologies.
© 2026 Dilate Technologies