The Data Tools initiative is the Data Tools-tagged component of the Innovation & Industry area of the U.S. National Science Foundation Center for Molecularly Optimized Networks (NSF MONET). It develops open-source notations, data schemas, and software platforms that make complex polymer-network information searchable, interoperable, and reusable across research and application settings.
Why this matters
Polymer-network data are difficult to compare because the materials combine stochastic molecular structures, network topology, processing history, and measured properties. The Center's Year 4 reporting narrative describes this as a practical barrier to using modern data science: researchers need ways to name, find, compare, and reuse data from different experiments and simulations.
Data Tools addresses that barrier through three connected layers:
BigSMILES provides a machine-readable line notation for stochastic polymer structures, including extensions for coarse-grained and non-covalent systems.
PolyDAT provides a standardized schema for polymer-characterization data and supports interoperability across tools and laboratories.
CRIPT carries these capabilities into a searchable, open-source polymer data platform linking synthesis, processing, and property measurements. CRIPT completed incorporation as a public benefit corporation in spring 2025.
Together, these tools lower the friction of sharing and reusing polymer data. That matters for faster materials discovery, more reproducible computational work, and translation of Center research into software and data infrastructure that can be used beyond the originating laboratories.
Documented reach and activity
Activity or outcome | Period | Documented result |
BigSMILES and polymer machine-learning benchmarking | 2024 study, reported in the Year 4 report | SMILES and BigSMILES were assessed across 12 polymer machine-learning tasks; the report states that BigSMILES enabled 24% faster training in large-language-model frameworks than competing notations while maintaining predictive accuracy for homopolymers and improving chemical encoding for copolymers. |
Validated BigSMILES conversion workflow | 2024 | A published workflow validated automated conversion from SMILES to BigSMILES for homopolymeric macromolecules. The Center's impact tracker attributes 4.9 million automatically generated BigSMILES records to this work. |
CRIPT incorporation | Spring 2025 | CRIPT completed incorporation as a public benefit corporation and was described in the Year 4 report as a searchable, open-source polymer database linking synthesis, processing, and property measurements. |
CRIPT soft launch | Mid-2025 | The platform began generating revenue through data-entry services and subscriptions. The recorded subscription prices were 120 USD/year for non-commercial users and 1,200 USD/year for commercial users; the revenue amount was not disclosed pending board approval. |
External use tracked across the polymer community | Updated November 12, 2025 | The Center's tracking record counted 273 papers citing CRIPT, BigSMILES, or PolyDAT; 247 (90.5%) were attributed to non-MONET-affiliated authors. The same record identified seven external research teams with documented active use of the tools. These are database-tracking figures, not an independent audit. |
Evidence that the approach works
Measured computational benefit
The Year 4 report identifies a 2024 benchmark across 12 polymer machine-learning tasks as evidence that BigSMILES can support large-scale computational work without sacrificing predictive accuracy for homopolymers. The reported 24% training-time advantage in large-language-model frameworks is a measured comparison, not an attendance or self-report metric. The Center's tracking record also summarizes separate reported findings of reduced prediction error and faster preprocessing; those figures are kept distinct from the Year 4 benchmark because they come from a different study and source trail.
Reusable data infrastructure
The published BigSMILES conversion workflow demonstrates a concrete interoperability function: it translates conventional SMILES representations into validated BigSMILES representations that can be reused in polymer-informatics datasets. The associated record describes the workflow as reducing translation inconsistencies and supporting dataset reuse. The Year 4 report further describes BigSMILES-to-structure drawing software, polymer search, and similarity-scoring tools being transitioned to the broader user community through CRIPT.
External validation and adoption
The Center's external-use tracker applies a defined inclusion criterion: documented use in a publication, conference proceeding, technical report, or patent filing; citation or use of BigSMILES, PolyDAT, or CRIPT; and authorship by an institution not directly affiliated with NSF MONET. Its November 2025 update lists examples including automated BigSMILES conversion, BigSMILES benchmarking, CRIPT-based data management, and PolyDAT schemas adapted for solid-electrolyte research. The tracker is useful evidence of reach, while the page preserves the distinction between internally recorded adoption and independently verified impact.
An independent 2024 Macromolecules perspective identifies the lack of open standards for exchanging polymer hardware, software, and data as a barrier to automation and machine learning, and discusses PolyDAT and BigSMILES in that context. The corresponding database record explicitly notes that the article does not provide concrete examples attributing utility or impact to NSF MONET tools; it is therefore used here as field context, not as a quantified outcome.
Response to reviewer feedback
Review feedback identified a need for simple software for generating BigSMILES descriptors, clear training materials, and practical examples. The documented response identifies the BigSMILES Builder, the PolyDAT Data Form, integrations with free chemical-drawing software, and video tutorials as existing or expanding resources. The feedback record marks the general response as implemented, while a related publication-focused implementation remains in progress.
What comes next
The Center's current tracking record sets a goal of at least 25 external research laboratories actively using and publishing with the tools; its November 2025 update documented seven external research teams. The documented next steps are to continue quarterly review of cited papers, expand training materials and practical examples, encourage use of BigSMILES in Center publications, and explore further integration with chemical-drawing software and commercial vendors.
For CRIPT, the 2025 soft-launch record describes additional features in the pipeline and a possible progression from individual subscriptions toward institutional licenses. Those plans are reported as future work rather than completed outcomes.







