Changelog¶
All notable changes to gs2txt will be documented in this page.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
Changelog¶
All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
Unreleased¶
0.1.0 - 2025-01-13¶
Added¶
- Core Features
GeneSetAnnotator- Main annotation engine for gene set analysis- Built-in gene filtering by p-value and log2FC thresholds
- Automatic pathway enrichment using GSEApy
-
Customizable prompt templates for domain-specific use cases
-
LLM Providers
OpenAIProvider- OpenAI API integration (GPT-4, GPT-3.5)AnthropicProvider- Anthropic Claude integration-
LiteLLMProvider- Unified interface for multiple LLM backends -
Enrichment Methods
PathwayEnrichment- GSEApy-based enrichment (MSigDB, KEGG, GO)-
CustomEnrichment- Adapter for pre-computed results -
Batch Processing
BatchProcessor- Process multiple gene sets from CSV files- Support for grouped data (e.g., by cluster)
-
Automatic output formatting (gs, annotation columns)
-
Two-Stage Pipeline (New in 0.1.0)
TwoStagePipeline.preprocess()- Single file preprocessingTwoStagePipeline.preprocess_batch()- Batch preprocessing from foldersTwoStagePipeline.annotate()- LLM annotation stage- YAML configuration support
-
Multiple pathway source merging (GO, KEGG, Reactome)
-
CLI Interface
gs2txtcommand for processing CSV files- Support for multiple providers and models
-
Configurable filtering parameters
-
Documentation
- Three usage scenarios with example scripts
- Comprehensive README with API examples
- MkDocs documentation site
Example Scripts¶
scenario1_batch_csv.py- Batch CSV processingscenario2_api.py- Python API integrationscenario3_two_stage.py- Two-stage pipeline processing