Clinical & Molecular Biomedicine

Gold Open AccessISSN pending
Translating Research into Better Healthcare

About Article

Research article • Volume 1, Issue 1, Pages 1-9 (2026) 100001

Genome-wide annotation and structural modeling of hypothetical proteins in Listeria monocytogenes

PDF
Check for updates

Abstract

The rapid expansion of genome sequencing projects has resulted in the identification of numerous hypothetical proteins whose functions remain uncharacterized. In Listeria monocytogenes serotype 4b, a major food-borne pathogen associated with high mortality rates, several predicted proteins lack functional annotation despite their potential role in pathogenicity and survival. In the present study, a comprehensive in silico approach was employed to functionally annotate 92 hypothetical proteins identified from the genome of L. monocytogenes. Protein sequences were retrieved and analyzed using sequence similarity searches, conserved domain identification, motif analysis, and multiple sequence alignment. Functional classification was performed based on BLAST, Pfam, and InterPro analyses. Structural prediction was performed via homology modeling via the SWISS-MODEL server, followed by structural validation and comparative analysis using PyMOL and DALI-Lite. Functional inference from structural models was supported by conserved-residue mapping and ProFunc analysis. Sequence-based analysis enabled classification of most hypothetical proteins into functional groups, including hydrolases, transferases, transporters, kinases, stress-response proteins, membrane proteins, DNA-binding proteins, and ATP-binding proteins. Structure-based modeling of selected proteins further confirmed the predicted catalytic residues, metal-binding sites, ligand-interaction sites, and conserved functional motifs. Several proteins were predicted to be involved in enzymatic activity, nucleotide metabolism, membrane transport, transcriptional regulation, and stress adaptation. This integrative computational analysis provides functional insights into previously uncharacterized proteins of L. monocytogenes. The findings enhance genome annotation quality and identify potential targets for further experimental validation, contributing to a better understanding of bacterial physiology and pathogenesis.

Keywords: Listeria monocytogenes, hypothetical proteins, functional annotation, homology modeling, genome analysis, structural prediction, protein function prediction

Introduction

Due to the speed and cost-effectiveness of genome sequencing, many small bacterial, archaeal, and eukaryotic genomes have been sequenced, and larger eukaryotic genomes are expected to be fully sequenced soon. In such a case, annotation becomes problematic when genes are predicted faster than annotation can keep pace. Despite ongoing improvements in genome annotation, a substantial proportion of open reading frames remain classified as "conserved hypothetical proteins" (Galperin and Koonin, 2004). Hypothetical proteins are predicted from nucleotide sequences but lacking direct experimental validation at the protein level. In many instances, these predicted proteins show limited sequence similarity to previously characterized and annotated proteins. Conserved hypothetical proteins constitute a significant fraction of genes in sequenced genomes; they encode proteins identified across multiple phylogenetic lineages but remain uncharacterized with respect to biological function or biochemical properties, sometimes accounting for a large portion of the predicted proteome (Sivashankari and Shanmughavel, 2006).

Determining the functions of hypothetical proteins is essential for improving genomic and proteomic annotations. The identification of new hypothetical proteins can reveal previously unknown structural folds and biological activities (Lubec et al., 2005). As more structures are characterized, novel domains, motifs, and conformational arrangements are likely to emerge, contributing to a clearer understanding of structure-function relationships in proteins. Functional characterization may also uncover new molecular pathways and regulatory cascades. Furthermore, hypothetical proteins may serve as biomarkers or therapeutic targets. Their identification can facilitate the discovery of previously unrecognized or computationally predicted genes. Ultimately, elucidating the roles of these proteins will deepen our understanding of biological systems and may contribute to the development of improved therapeutic interventions.

The biological roles of many predicted proteins remain unknown because their existence has been inferred solely from genomic sequences. The rapid completion of genome sequencing projects has generated extensive biological datasets, enabling the association of genes with their corresponding protein products, which are fundamental to cellular function. Advances in genome and cDNA sequencing technologies allow the prediction of numerous proteins for which experimental evidence is lacking. Functional insights into hypothetical proteins are often derived from sequence comparisons with proteins of known function in model organisms. Since similarities in sequence or conserved domains frequently indicate shared functional properties, such analyses can suggest potential biological roles. Moreover, if a hypothetical protein exhibits significant structural homology with characterized proteins, it is reasonable to infer comparable molecular features.

Traditional biochemical and molecular biology approaches can accurately assign functions to genes; however, these experimental methods are labor-intensive and costly. Consequently, even in fully sequenced genomes, only about 50-60% of genes have been functionally annotated. Automated genome analysis and computational annotation strategies provide alternative means to interpret genomic information. Therefore, determining protein function remains a major challenge in the post-genomic era. This has increased reliance on bioinformatics tools to predict the roles of uncharacterized protein sequences, and several contemporary computational approaches have been developed to address this issue.

Listeria monocytogenes is the causative agent of listeriosis. It is a facultative anaerobic bacterium capable of surviving under both aerobic and anaerobic conditions. The organism can invade and proliferate within host cells and is recognized as one of the most severe foodborne pathogens, with mortality rates of 20-30% in clinical cases (Gebremedhin et al., 2021). In the United States, listeriosis is responsible for approximately 2,500 reported cases and 500 deaths annually, making it one of the leading causes of death among foodborne bacterial infections. Although relatively rare, the disease is associated with a high case-fatality rate. Its emergence as a public health concern has been attributed to changing dietary habits, advances in food preservation technologies that extend shelf life, and the bacterium's ability to survive and grow at refrigeration temperatures.

Disease caused by Listeria monocytogenes is called Listeriosis (a bacterial infection). In most cases, these infections occur in immunocompromised individuals, such as pregnant women (they are about 20 times more likely to get infected than healthy adults), Newborns, persons with diabetes, cancer, and kidney disease, persons with AIDS, and persons who take glucocorticosteroid medications. Listeriosis is fatal in at least 1 in 5 infected individuals. Listeria monocytogenes causes flu-like symptoms and meningitis if it spreads to the central nervous system. The first reported case of Listeriosis was in 1981; since then, numerous listeriosis outbreaks have been detected and investigated. Subsequent studies have confirmed that the Listeria monocytogenes serovar primarily responsible for Listeria infection is serotype 4b (Nuesch-Inderbinen et al., 2021). So, of the 13 serovars of Listeria monocytogens, Listeria monocytogenes 4b is the most prominent one. Listeria monocytogenes has been reported to cause miscarriages, stillbirths, or even premature pregnancies among pregnant women with underlying diseases. The chances of a fetus being affected by an affected mother are high.

The Listeria monocytogenes genome is approximately 3 Mb. Its genome consists of one chromosome. Total number of nucleotides is 2912690, total number of protein genes found to be 2766, and total number of RNA genes is 85. Many stress proteins are expressed in Listeria monocytogenes under stressful conditions, helping to bacterium survive harsh environments for extended periods. There are several important facts about Listeria monocytogenes genome available. Genome analysis suggests that the Listeria monocytogenes 4 b genome contains several hypothetical proteins. There are 92 hypothetical proteins in Listeria monocytogenes 4 b. There is little information available on these hypothetical proteins from Listeria monocytogenes genome. Here, our aim is to annotate the hypothetical proteins from Listeria monocytogenes by using recent bioinformatics tools. After annotating a genome, we grouped genes into two types: the first comprises known genes with functional characterization; the second comprises conserved hypothetical genes conserved across many organisms. Generally, a newly sequenced genome contains approximately 30% of genes that are poorly or incorrectly annotated. Hence, these poorly characterized hypothetical genes might play an important role in our understanding of the pathogen's pathogenesis and virulence. Given that so little is known about these genes, we have attempted to assign plausible functions to hypothetical proteins. The functional evidence suggests that these hypothetical proteins may have specific functions, but further detailed analysis to needed to confirm their functions.

2. Materials and Methods

2.1. Data mining

We have analyzed the Listeria monocytogenes's genome and the hypothetical proteins that it contains by retrieving the information at the P1R website. We have identified and downloaded 92 hypothetical protein sequences of Listeria monocytogenes. The sequences of all proteins analyzed with their primary accession number, length, molecular mass, and proposed functions (Table 1). Protein sequences are saved in FASTA format for further processing. A systematic pipeline used in this study is illustrated in Figure 1.

Figure 1. Outline of structure and function assignment to the hypothetical proteins.

Download: Download full-size image

Table 1. List of hypothetical proteins in Listeria monocytogenes.

S. No.Name of Hypothetical ProteinpILengthPfamProposed Function
1C1L300/ C1L300_LISMC9.69447MatE(PF01554) antiporter activity, drug transmembrane transporter activity
2C1L2X5/ C1L2X5_LISMC7.85345UPF0118(PF01594) NA
3C1L2V8/ C1L2V8_LISMC5.02120DUF964(PF06133) NA
4C1L2V7/ C1L2V7_LISMC5.45267YmdB(PF13277) putative phosphoesterases
5C1L2S2/ C1L2S2_LISMC6.41274FtsJ(PF01728) S4 (PF01479) methyltransferase involved in viral RNA capping, RNA
6C1L2Q9/ C1L2Q9_LISMC6.47321DUF1385(PF07136) NA
7C1L2P0/ C1L2P0_LISMC4.78118No domainsNA
8C1L2N2/ C1L2N2_LISMC5.4492DUF503(PF04456) NA
9C1L2K8/ C1L2K8_LISMC5.38174Acetyltransf_3(PF13302) transfers acetyl group
10C1L2J0/ C1L2J0_LISMC7.84103NANA
11C1L2E5/ C1L2E5_LISMC5.13174Metallophos_2(PF12850) Phosphoesterase (hydrolase activity, acting on ester bonds)
12C1L1Z7/ C1L1Z7_LISMC6.42774Glycos_transf_2Glycos_transf_2 (PF00535); Glyphos_transf (PF04464)
13C1L1V8/ C1L1V8_LISMC9.43267DUF817(PF05675) NA
14C1L1U8/ C1L1U8_LISMC6.18555Lactamase_BLactamase_B (PF00753); RNA binding; hydrolase activity, acting on ester bonds, metal-dependent
15C1L1S5/ C1L1S5_LISMC10.1344DUF218(PF02698) NA
16C1L1R3/ C1L1R3_LISMC9.3255TerC(PF03741) integral to membrane
17C1L1R2/ C1L1R2_LISMC9.36255TerC(PF03741) integral to membrane
18C1L1R1/ C1L1R1_LISMC9.09453MatE(PF01554) antiporter activity, drug transmembrane transporter activity
19C1L1K4/ C1L1K4_LISMC9.81201SNARE_assoc(PF09335) NA
20C1L1G8/ C1L1G8_LISMC5.59725Tex_N(PF09371), HHH_3 (PF12836), S1 RNA binding, hydrolase activity- acting on ester bonds
21C1L1B3/ C1L1B3_LISMC4.78125Ribonuc_L-PSP(PF01042) Inhibits protein synthesis by cleaving the mRNA
22C1L190/ C1L190_LISMC5.09282NANA
23C1L165/ C1L165_LISMC4.69176YceI(PF04264) binds to polyisoprenoid
24C1L164/ C1L164_LISMC9.54306EamA(PF00892) transporter activity
25C1L162/ C1L162_LISMC9.27233DUF554(PF04474) NA
26C1L161/ C1L161_LISMC4.79296CN_hydrolase(PF00795) hydrolase activity, acting on carbon-nitrogen (but not peptide)
27C1L158/ C1L158_LISMC5.68282PhzC-PhzF(PF02567) Phenazine biosynthes
28C1L143/ C1L143_LISMC5.33309DAGK_cat(PF00781) NAD+ kinase activity, diacylglycerol kinase activity
29C1L0U3/ C1L0U3_LISMC4.83288Hydrolase_3(PF08282) hydrolase activity
30C1L0T5/ C1L0T5_LISMC5.42209HhH-GPD(PF00730) DNA binding, endonuclease activity
31C1L0T0/ C1L0T0_LISMC9.82187DUF420(PF04238) NA
32C1L0R8/ C1L0R8_LISMC5.34606Sulfatase(PF00884) sulfuric ester hydrolase activity
33C1L0Q0/ C1L0Q0_LISMC9.71246TauE(PF01925) NA
34C1L0P3/ C1L0P3_LISMC4.87172Acetyltransf_1(PF00583) acetyltransferase activity
35C1L0M8/ C1L0M8_LISMC4.78110PadR(PF03551) involved in negative regulation of phenolic acid metabolism
36C1L0L0/ C1L0L0_LISMC6.08394Methyltrans_SAM(PF10672) methyltransferase activity
37C1L0K2/ C1L0K2_LISMC7.7431Xan_ur_permease(PF00860) transporter activity
38C1L0G9/ C1L0G9_LISMC4.82264AP_endonuc_2(PF01261) involved in the myo-inositol catabolism (isomerase activity)
39C1L075/ C1L075_LISMC5.18346Lactonase(PF10282) 6-phosphogluconolactonase activity
40C1L028/ C1L028_LISMC5.72143Usp(PF00582) response to stress
41C1L024/ C1L024_LISMC5.26226GATase(PF00117) transferase activity
42C1KZV6/ C1KZV6_LISMC4.89281NmrA(PF05368) nucleotide binding (negative transcriptional regulator
43C1KZU1/ C1KZU1_LISMC5.41377Gly_kinase(PF02595) glycerate kinase activity
44C1KZQ3/ C1KZQ3_LISMC4.64279Hydrolase_3(PF08282) hydrolase activity
45C1KZM7/ C1KZM7_LISMC8.93156Usp(PF00582) response to stress
46C1KZM4/ C1KZM4_LISMC6.15120YjbR(PF04237) NA
47C1KZE1/ C1KZE1_LISMC4.83270Hydrolase_3(PF08282) hydrolase activity
48C1KZ86/ C1KZ86_LISMC5.18111PhnA_Zn_Ribbon(PF08274), PhnA NA
49C1KZ81/ C1KZ81_LISMC5.97494FTR1(PF03239) transmembrane transport
50C1KZ17/ C1KZ17_LISMC8.78220MgtC(PF02308) transport of Mg2+
51C1KZ02/ C1KZ02_LISMC5.31280SAM_adeno_trans(PF01887) NA
52C1KYZ6/ C1KYZ6_LISMC8.7362MacB_PCD(PF12704), FtsX (PF02687) transport lipids targeted to the outer membrane across the
53C1KYZ4/ C1KYZ4_LISMC5.4497ABM(PF03992) monooxygenase activity
54C1KYZ2/ C1KYZ2_LISMC4.62275Hydrolase_3(PF08282) hydrolase activity
55C1KYY1/ C1KYY1_LISMC5.4440HD(PF01966) metal ion binding, phosphoric diester hydrolase activity
56C1KYX9/ C1KYX9_LISMC6.83217Peptidase_M50(PF02163) proteolysis, metalloendopeptidase activity
57C1KYX3/ C1KYX3_LISMC9.03306DAGK_cat(PF00781) kinase, transferase
58C1KYW9/ C1KYW9_LISMC10.1357UPF0104(PF03706) NA
59C1KYT3/ C1KYT3_LISMC5.78211UPF0029(PF01205), DUF1949 (PF09186) GTP binding
60C1KYS5/ C1KYS5_LISMC6.43287DUF161(PF02588), DUF2179 (PF10035) NA
61C1KYP0/ C1KYP0_LISMC4.97322UPF0052(PF01933) NA
62C1KYM2/ C1KYM2_LISMC5.12259CN_hydrolase(PF00795) hydrolase activity, acting on carbon-nitrogen (but not peptide)
63C1KYL6/ C1KYL6_LISMC5.39273Hydrolase_3(PF08282) ATP binding, ATPase activity, coupled to transmembrane
64C1KYI4/ C1KYI4_LISMC5.98201MTS(PF05175) 16S rRNA (guanine(966)-N(2))-methyltransferase activity
65C1KY64/ C1KY64_LISMC5.76117ArsC(PF03960) regulate the transcription of multiple genes in response to
66C1KY61/ C1KY61_LISMC5.27291Cation_efflux(PF01545) cation transmembrane transporter activity
67C1KY51/ C1KY51_LISMC5.94147NifU_N(PF01592) iron ion binding, iron-sulfur cluster binding
68C1KY50/ C1KY50_LISMC4.87464UPF0051(PF01458) iron-sulfur cluster assembly
69C1KY46/ C1KY46_LISMC9.02279TauE(PF01925) involved in the transport of anions across the cytoplasmic
70C1KY45/ C1KY45_LISMC5.08462Metallophos(PF00149), 5_nucleotid_C hydrolase activity (5_nucleotid_C)
71C1KY43/ C1KY43_LISMC4.94255Hydrolase_like(PF13242), Hydrolase_6 hydrolase activity
72C1KY41/ C1KY41_LISMC4.82436DUF21(PF01595), CBS (PF00571), flavin adenine dinucleotide binding (CBS),
73C1KY30/ C1KY30_LISMC9.43406Voltage_CLC(PF00654) voltage-gated chloride channel activity
74C1KY03/ C1KY03_LISMC5.29156Rrf2(PF02082) NA
75C1KXZ5/ C1KXZ5_LISMC4.65281Hydrolase_3(PF08282) hydrolase activity
76C1KXZ0/ C1KXZ0_LISMC4.76532Amidohydro_3(PF07969) hydrolase activity, acting on carbon-nitrogen (but not peptide)
77C1KXY4/ C1KXY4_LISMC6.26257Methyltransf_26(PF13659) methyltransferase activity
78C1KXX7/ C1KXX7_LISMC4.83270Hydrolase_3(PF08282) hydrolase activity
79C1KXV5/ C1KXV5_LISMC5.52249EAL(PF00563) NA
80C1KXN7/ C1KXN7_LISMC6.53331Bac_luciferase(PF00296) oxidoreductase activity, acting on paired donors, with
81C1KXN1/ C1KXN1_LISMC6.58147DUF523(PF04463) NA
82C1KWJ0/ C1KWJ0_LISMC6.74121DUF1798(PF08807) NA
83C1KWG8/ C1KWG8_LISMC6.52327DUF939(PF06081), DUF939_C NA
84C1KWG7/ C1KWG7_LISMC4.39125Glyoxalase_2(PF12681) NA
85C1KWG4/ C1KWG4_LISMC6.52205HTH_11(PF08279), CBS (PF00571) sequence-specific DNA binding transcription factor activity
86C1KWC7/ C1KWC7_LISMC5.29291YicC_N(PF03755), DUF1732 (PF08340) play a role in the stationary phase survival (YicC),
87C1KW03/ C1KW03_LISMC6.11126OsmC(PF02566) has a novel pattern of oxidative stress regulation
88C1KVW1/ C1KVW1_LISMC6.22192rRNA_methylase(PF06962) methyltransferase activity
89C1KVW0/ C1KVW0_LISMC5.48321Radical_SAM(PF04055) catalytic activity, iron-sulfur cluster binding
90C1KVA0/ C1KVA0_LISMC6.02234DUF633(PF04816) tRNA (adenine-N1-)-methyltransferase activity
91C1KV99/ C1KV99_LISMC5.5373NIF3(PF01784) NIF3 interacts with the yeast transcriptional coactivator
92C1KV71/ C1KV71_LISMC4.92269Hydrolase_3(PF08282) cation transport, ATP binding, ATPase activity, coupled to

2.3. Structure Prediction

The three-dimensional structures of hypothetical proteins were examined to infer their potential biological functions. Structural prediction was carried out using homology modeling, which generally involves multiple sequential steps. First, a suitable template protein with an experimentally determined structure, typically belonging to the same or a closely related family, is identified. Second, appropriate sequence alignment tools are employed to obtain the best possible alignment between the target and template proteins. Third, the target sequence is divided into smaller segments to facilitate accurate modeling. Fourth, corresponding fragments are searched in structural databases based on sequence similarity and the spatial conformation of the selected template. Fifth, the spatial coordinates of these matched segments are assembled and adjusted to construct the full-length target structure, ensuring that all atomic positions are properly incorporated. These procedures are repeated multiple times to generate several models, from which an averaged structure is obtained, followed by comprehensive energy minimization to refine the final predicted model. The 3-D structures were retrieved using the SWISS-MODEL server, an automated comparative modeling system for predicting protein 3-D structures. The entire homology modeling workflow can be performed in the 'project mode' of PyMOL, an integrated sequence-structure workbench. Using the SWISS-MODEL server, we obtained suitable 3D structures for thirteen hypothetical proteins selected for detailed structural analysis, and the resulting structural models were downloaded in PDB format for further analysis.

2.4. Function prediction

Clustering of gene expression profiles is widely used for functional prediction, based on the principle that genes involved in related biological processes tend to exhibit coordinated expression patterns. Functional inference can also be derived from protein-protein interaction networks, where the role of an uncharacterized protein is estimated from the known functions of its interacting partners. Schwikowski and co-workers introduced a neighbor-counting strategy in which the probable function of an unknown protein is determined according to the frequency of functional annotations among its interaction partners. This strategy was later improved by Ishigaki et al., 2001 through the incorporation of a chi-square (^2) statistical framework to enhance prediction reliability. In both methods, all functional contributions from neighboring proteins are treated with equal significance. Several publicly accessible signature and motif databases support functional annotation and are frequently integrated with sequence clustering and domain-based resources. These include PROSITE, PRINTS, Pfam, ProDom, Blocks, SMART, and InterPro; in the present study, Pfam and InterPro were specifically utilized. Structural classification of proteins into homologous families, superfamilies, and folds can be performed using databases such as SCOP, CATH, FSSP, CAMPASS, and HOMSTRAD. Additionally, the ProFunc server facilitates structure-based functional prediction by comparing modeled protein structures against multiple databases and reporting significant matches for functional interpretation.

2.5. Visualization

We carried out structural analysis using PyMOL for superposition of molecules, secondary-structure-based alignment, and calculation of RMSD values. PyMOL is also used to depict structures, as well as to analyze potential residue positions and functions. We visualized the template residues and the model after superimposing the model onto the template. Each residue was visualized individually, and residue conservation was assessed.

2.6. Structure analysis

Structure-based comparison approaches are highly effective for detecting remote evolutionary relationships that may not be apparent from sequence similarity alone. Structural information also enables the recognition of functional sites that have arisen independently during evolution. Because protein function is inherently dependent on three-dimensional conformation, structural analysis provides direct insight into the molecular mechanisms underlying biological activity. Local structural alignment focuses on identifying conserved three-dimensional arrangements of specific residues, even among proteins that differ substantially in overall fold. Several computational tools are available for such analyses, including DaliLite, which utilizes C-C distance matrix comparisons; TOPS; SSM; VAST; GRATH, which represents secondary structure elements as graph nodes connected by spatial relationships; and SSAP, which evaluates structural similarity based on C comparisons and additional scoring features. Other methods such as LSQMAN, CE (Combinatorial Extension), MATRAS, and LOCK2 also employ atom-based or secondary structure vector representations for alignment. In the present study, DALI-Light was selected for structural comparison and analysis.

3. Results and Discussion

3.1. Sequence-based Function Prediction

Information derived from BLAST, ClustalW, and other online servers for sequence analysis revealed close similarities of hypothetical proteins to proteins of known function. Based on the proposed functions, we classified all 92 hypothetical proteins into functional classes. All the Hypothetical proteins have been classified into the following categories.

Hydrolase. The genomic proteins of Listeria monocytogenes most abundantly showed similarity with proteins having hydrolase function, and hence, by comparing the conserved residues, we can assume the proteins to have the same function as well. The sequence of C1L0U3, C1L0R8, C1KZQ3, C1KZE1, C1KYZ2, C1KYY1, C1KY45, C1KY43, and C1KXZ5 has been predicted to have hydrolase function following a precise analysis that compared its sequence similarity with that of known, characterized proteins. Hydrolases are enzymes that facilitate hydrolytic reactions, in which a water molecule is utilized to break chemical bonds within a substrate. During this process, the hydrogen (H^+) and hydroxyl (OH^-) components of water are incorporated into the reacting molecule, resulting in its cleavage into two or more smaller products. The term "hydrolase" serves as the systematic designation for enzymes classified under Enzyme Commission (EC) class 3, which encompasses all enzymes that catalyze hydrolysis reactions.

For the C1L0U3 gene, 65 matching sequences were found by FASTA search. 8 neighboring genes found as Q8Y970, D2PAB8, D2NZC4, C1L0U3, E1UE28, B8DIC3, Q92DZ3, D3UKJ4 having the percentage identity of 100%

Transferase. Transferases are enzymes that catalyze the movement of specific chemical groups, such as methyl, glycosyl, acyl, or phosphate groups, from one molecule to another. In these reactions, one compound acts as the donor of the functional group, while another serves as the acceptor. The nomenclature and classification of transferases follow the format "donor:acceptor group transferase," reflecting the nature of the transferred group and the participating substrates. Proteins belonging to this group are C1L2K8, C1L1Z7, C1L0P3, C1L0L0, C1L024, C1KYI4, C1KXY4, C1KVW1, and C1KVA0. The protein sequence has been found to be similar to that of characterized proteins in this category. This group is the second most abundant in Listeria monocytogenes, after hydrolases. For the gene C1L0L0, 65 matching sequences were identified by a FASTA search, and 8 motifs were matched by InterProScan, which searches the InterPro database. The 7 neighbouring genes are as follows D2PA32, D2NYL3, Q8Y9E8, E1UDU6, B8DA44, Q92E70 and D3USH9 with percentage identity as 99.5%

Transporter. Transporter proteins are proteins that move substances within an organism. Transport proteins are vital to the growth and life of all living things. There are several different kinds of transport proteins. The proteins found in Listeria monocytogenes included in this category are C1L300, C1L1R1, C1L164, C1L0K2, C1KZ81, C1KZ17, C1KYZ6, C1KY61, and C1KY46. For the C1L1R1 gene, 4 matching sequences were found by FASTA search, and 6 motifs were found by InterPro scan. There were 7 neighboring genes found which are E1UF08, B8DEE5, Q8Y8B8, D2P3J7, D2P0R1, Q92D31, and D3ULQ9 with percentage identity as follows 97.9%

RNA Binding. The following residues belong to this group: C1L2S2, C1L1U8, and C1L1G8. RNA-binding proteins are proteins that bind to RNA recognition motifs (RRMs) of double or single-stranded RNA in cells and participate in forming ribonucleoprotein complexes. They are cytoplasmic and nuclear proteins. For the C1L2S2 gene, 15 matching sequences were identified by FASTA search, and 11 sequence motifs were identified in InterPro. The 7 neighbouring genes are Q8Y7C0, D2P4S6, D2P1Z0, E1UG75, B8DFW2, Q92BY9 and D3UMV0 with sequence identity as 97.4%

Hydrolase acting on carbon-nitrogen but not on peptide. Following mentioned are the proteins C1L161, C1KYM2, and C1KXZ0, which belong to this category. This family contains hydrolases that break carbon-nitrogen bonds. This sub family of hydrolase is numbered as EC 3.5. There were 35 matching sequences found by FASTA search and 6 sequence motifs matched in the InterPro scan. There were 6 neighbouring genes found which are E1UD07, D2P8J8, D2NWP2, B8DEW3, Q8YA77, and Q92EZ8 with percentage identity as 99.2%

Kinase. Kinases are enzymes that catalyze the transfer of phosphate groups from high-energy donor molecules, typically adenosine triphosphate (ATP), to specific target substrates in a reaction known as phosphorylation. Through this process, kinases regulate numerous cellular activities, including signal transduction, metabolism, and cell cycle progression. They belong to the broader class of phosphotransferases. In the present study, proteins C1L143, C1KZU1, and C1KYX3 were identified as members of this enzyme category.

Response to stress. Cells exposed to stressful conditions often accumulate abnormal, misfolded, or damaged proteins as a consequence of environmental or physiological challenges. To mitigate these effects, cells increase the production of stress-related proteins that assist in stabilizing, refolding, or degrading altered proteins. Rather than solely enabling survival under extreme or lethal stress, these stress proteins also play a crucial role in cellular recovery following stress exposure, thereby helping restore normal cellular homeostasis. C1LO28, C1KZM7, and C1KW03 are the proteins that belong to this category. A FASTA search of the C1KZM7 gene identified 21 matching sequences, and an InterPro scan identified 4 sequence motifs. Q927G8, E1UBQ4, D2P7S2, D2NVX3, B8DAW9, Q8Y405 and D3USS9 are the 7 neighbouring genes found with 99.3%

Integral to membrane. An integral membrane protein is a protein molecule, or protein complex, that is permanently embedded within the plasma membrane through hydrophobic regions that interact strongly with the lipid bilayer. These proteins are tightly associated with membrane phospholipids and cannot be easily removed without disrupting the membrane structure. Integral membrane proteins are broadly divided into two principal categories: (1) transmembrane proteins and (2) integral monotropic proteins. Transmembrane proteins represent the more prevalent group and extend completely across the lipid bilayer. Such proteins perform diverse biological functions, including acting as transporters, ion channels, receptors, enzymes, structural anchors, components involved in energy storage and signal transduction, and mediators of cell-cell adhesion. Representative examples include integrins, cadherins, the insulin receptor, neural cell adhesion molecules (NCAM), selectins, glycophorin, and rhodopsin. C1L1R3, C1L1R2, and C1KY64 are the proteins that belong to this group. For the gene C1KY64, we found 39 matching sequences by FASTA search and 5 motifs matched in InterPro scan. 6 neighbouring genes found were E1UAU6, B8DDD2, Q8Y4L1, D2P980, D2NY93, and Q928L2 with percentage identity as 98.3%

DNA Binding. DNA-binding proteins are proteins that contain a DNA-binding domain, which in turn is responsible for binding to single or double-stranded DNA molecules. C1L0T5 and C1KWG4 are the proteins that have this domain and are therefore classified in this category.

Phosphoesterase. Phosphoesterase is a family of hydrolase with acts on ester bonds. Phosphoesterases are involved in cleaving ester bonds. Proteins belonging to this category are C1L2V7 and C1L2E5. For gene C1L2E5 there are 24 matching sequences found by FASTA search and 6 sequence motifs matched in InterPro. There are 7 neighbouring genes which are Q8Y7N4, E1UFM2, B8DHZ2, D2P485, D2P1E9, Q92CG9 and D3UME0 with percentage identity as 100%

ATP Binding. An ATP-binding protein is a protein that specifically interacts with adenosine 5′-triphosphate (ATP), a ribonucleotide composed of the purine base adenine attached to the sugar D-ribofuranose and linked to three phosphate groups. ATP serves as a primary carrier of chemical energy and phosphate groups within the cell. Proteins that bind ATP typically utilize this interaction to drive biochemical reactions, regulate cellular processes, or facilitate molecular transport, thereby playing essential roles in energy-dependent cellular functions. C1KYL6 and C1KV71 are the proteins in this category. For the C1KYL6 gene, we identified 69 matching sequences via a FASTA search. D2P8J2, D2NWN6, Q8YA83, Q92F06, D3URG2, E1UCZ8, and B8DEX2 are the 7 neighboring genes found with percentage identities as 98.9%

Other Proteins. Proteins in this category are proteins that are one of their kind, i.e., no other protein in Listeria monocytogenes shares the same function as they do. Proteins like C1L1B3 whose function is mRNA cleavage, C1L165 whose function is to bind to polyisoprenoid, C1L158 which is responsible for phenazine biosynthesis, C1L0M8 which is involved in negative regulation of phenolic acid metabolism, C1K0G9 whose function is to catabolize myo-inositol, C1L075 which is polyphosphogluconolactonase, C1KZV6 which is responsible for nucleotide binding, C1KYZ4 which is monoxygenase, C1KYX9 which is involved in proteolysis, C1KYT3 is a GTP binding protein, C1KY51 is iron binding protein, C1KY50 has iron-sulphur cluster assembly, C1KY30 has chloride chemical activity, C1KXN7 is an oxidoreductase, C1KWC7 has a role in stationary phase survival, C1KVW8 has catalytic activity. The proteins are among those that don't share any function with any other protein. This is why they are being classified in this category. For the C1L165 gene, 11 matching sequences were identified by a FASTA search, and 4 sequence motifs were identified by an InterPro scan. There were 7 neighbouring genes found which are as follows Q8Y8U6, E1UEF3, D2PB35, D2NZR0, B8DGH0, Q92DM4 and D3UL61 with percentage identities of 99.4%

Unknown. Proteins included in this category are the proteins whose function could not be predicted due to a lack of any known characterized protein hits. Due to the unavailability of good hits, we classified them in this category, as no function can be assigned to them until we obtain hits of characterized proteins whose function is known to us, instead of hypothetical, uncharacterized, or putative proteins, which we obtained for proteins of this category.

The major functional classes identified in this analysis are summarized in Table 2.

Table 2. Major Classes of Hypothetical Proteins in Listeria monocytogenes.

S. No.Proposed function of Hypothetical proteinPrimary Accession number
1HydrolaseC1L0U3, C1L0R8, C1KZQ3, C1KZE1, C1KYZ2, C1KYY1, C1KY45, C1KY43, C1KXZ5
2TransferaseC1L2K8, C1L1Z7, C1L0P3, C1L0L0, C1L024, C1KYI4, C1KXY4, C1KVW1, C1KVA0
3TransporterC1L300, C1L1R1, C1L164, C1L0K2, C1KZ81, C1KZ17, C1KYZ6, C1KY61, C1KY46
4RNA bindingC1L2S2, C1L1U8, C1L1G8
5Hydrolase acting on carbon-nitrogen but not on peptideC1L161, C1KYM2, C1KXZ0
6KinaseC1L143, C1KZU1, C1KYX3
7Response to stressC1LO28, C1KZM7, C1KW03
8Integral membraneC1L1R3, C1L1R2, C1KY64
9DNA bindingC1L0T5, C1KWG4
10PhosphoesteraseC1L2V7, C1L2E5
11ATP BindingC1KYL6, C1KV71
12OthersC1L1B3, C1L165, C1L158, C1L0M8, C1L0G9, C1L075, C1KZV6, C1KYZ4, C1KYX9, C1KYT3, C1KY51, C1KY50, C1KY41, C1KY30, C1KXN7, C1KWC7, C1KVW8, C1KV99
13UnknownC1L2X5, C1L2V8, C1L2Q9, C1L2P0, C1L2J0, C1L1S5, C1L1K4, C1L198, C1L162, C1L0T0, C1L0Q0, C1KZM4, C1KZ86, C1KZ02, C1KYW9, C1KYS5, C1KYP0, C1KY03, C1KYN7, C1KXN1, C1KWJ0, C1KWG8, C1KWG7

3.2. Structure-based Function Prediction

We have sought to develop a template for building models of hypothetical proteins. Fortunately, we identified templates for 13 proteins and subsequently predicted their structures. To predict the function of hypothetical proteins, we have extensively analyzed the atomic coordinates of models and proposed corresponding functions based on structural similarity, fold, domain, motif, and analysis of structurally conserved residues.

C1L1U8. The protein C1L1U8 is 555 residues long, and the model was successfully built for residues 6-555. The overall structure of this protein is shown in Figure A. Our model shares 81.09%

C1KYM2. The protein C1KYM2 comprises 259 residues, and the model was successfully built for residues 2-259. The overall structure of this protein is shown in Figure B. Our model shares 37.31%

C1KYY1. The protein C1KYY1 comprises 440 residues, and its model was successfully built for residues 1-439. The overall structure of this protein is shown in Figure C. Our model shares 48.23%

C1L0M8. The protein C1L0M8 is 110 residues long, and the model was successfully built for residues 5-105. The overall structure of this protein is shown in Figure D. Our model shares 40.78%

C1L0R8. The protein C1L0R8 comprises 606 residues, and the model was successfully built for residues 201-605. The overall structure of this protein is shown in Figure E. Our model shares 30.07%

C1L165. The protein C1L165 is 176 residues long, and the model was successfully built for residues 5-176. The overall structure of this protein is shown in Figure F. Our model shares 41.28%

Figure 2. Structural superimposition of modeled proteins C1L1U8-C1L165 with their respective template structures. Cartoon representations showing the three-dimensional models superimposed over their corresponding template structures. (A) C1L1U8 (light green) superimposed on template 3ZQ4 (yellow). (B) C1KYM2 (deep purple) superimposed on template 2E11 (orange). (C) C1KYY1 (limon) superimposed on template 3IRH (deep teal). (D) C1L0M8 (blue) superimposed on template 4ESB (black). (E) C1L0R8 (red) superimposed on template 2W8D (brown). (F) C1L165 (green) superimposed on template 1WUB (dark blue).

Download: Download full-size image

Figure 3. Structural superimposition of modeled proteins C1L2E5-C1KY50 with their respective template structures. Cartoon representations depicting the structural alignment of modeled proteins with their corresponding templates. (A) C1L2E5 (red) superimposed on template 1S3L (green). (B) C1KZM7 (raspberry) superimposed on template 3HGM (wheat). (C) C1KYZ6 (red) superimposed on template 3FTJ (green). (D) C1KYT3 (green) superimposed on template 1VI7 (red). (E) C1KY51 (pink) superimposed on template 1SU0 (blue). (F) C1KYL6 (blue) superimposed on template 3R4C (pink). (G) C1KY50 (red) superimposed on template 2ZU0 (cyan).

Download: Download full-size image

Table 3. Superimposed residues between template and modeled proteins C1L1U8-C1L165.

Protein IDSuperimposed Residues (Template -> Model)
C1L1U8Asp78->Asp78, His79->His79, Asp164->Asp164, His390->His390, His74->His74, His76->His76, His142->His142, Asp195->Asp195, His368->His368, Gly49->Gly49, Asp51->Asp51, Asp443->Asp443, Glu464->Glu464, Asp449->Asp449, Arg544->Arg544, Ser366->Ser366, His364->His364, Gly367->Gly367, Ser233->Ser233, Glu77->Glu77, Phe42->Phe42, Tyr52->Tyr52
C1KYM2Glu43->Glu42, Lys109->Lys111, Cys143->Cys145, Tyr144->Tyr146, Phe49->Tyr48, Phe113->Phe115, Trp175->Trp171, Val142->Ile144, Pro176->Pro172
C1KYY1His129->His120, Glu122->Glu113, His66->His64, His110->His110, Asp111->Asp102, Asp183->Asp173, Lys14->Lys12, Asn36->Ala34, Gln41->Gln39, Arg326->Arg317, Lys330->Arg321, His114->His105, Tyr239->Tyr231, Arg63->Arg61, Leu49->Leu47, His119->His110, Tyr187->Tyr177, Tyr243->Tyr235, Tyr368->Tyr358
C1L0M8Gly25->Gly27, Tyr26->Tyr28, Glu42->Glu42, Arg71->Arg71, Lys72->Lys72, Tyr46->Tyr46, Arg51->Arg51
C1L0R8Thr297->Thr272, His412->His388, Glu253->Glu230, Trp350->Tyr325, His472->His449, Asp471->Asp448, His343->His318, Asn345->Asn320, Arg352->Arg327, Thr408->Ser384
C1L165His18->His21, Arg62->Arg66, His65->His69, Trp146->Tyr151

Table 4. Superimposed residues between template and modeled proteins C1L2E5-C1KY50.

Protein IDSuperimposed Residues (Template -> Model)
C1L2E5Asp8->Asp8, His10->His10, Asp36->Ser36, Asn59->Asn54, Asn60->Cys55, His97->His78, His120->His107, Thr121->Ser108, His122->His109
C1KZM7Arg135->Tyr130, Ser131->Ser126, Val38->Val40, Pro8->Gly10, Gly117->Gly112, Gln119->Thr114, Ala133->Ser128, Val132->Val127, Gly120->Gly115, Gly123->Ala118, Asn122->Ser117
C1KYZ6Leu469->Lys186, Tyr465->Trp182, Asn346->Ala79, Ile342->Thr75
C1KYT3Ser23->Ser21, His54->His52, Glu77->Glu74, Arg104->Arg100, Lys22->Lys20, Arg24->Arg22, Phe25->Phe23, Asp75->Asp71, Gly76->Gly72, Pro78->Pro74, Ala82->Ala78, Tyr105->Tyr101, Tyr106->Phe102, Gly107->Gly103, Leu111->Leu107, Leu116->Leu112, Tyr120->Tyr116, Asp74->Asp70, Thr81->Thr77
C1KY51Cys40->Cys41, Asp42->Asp43, Cys65->Cys66, Arg124->Arg125, Cys127->Cys128, Gly64->Gly65, Asp57->Ala58, Asp77->Gln78, Glu56->Val57
C1KYL6Trp171->Ser187, Phe175->Asn191, Arg45->Arg46, Asp1->Asp12, Thr43->Thr44, Asp8->Asp10, Lys185->Lys201, Asn211->Asn227, Asp208->Asp22
C1KY50Phe373->Phe415, Leu375->Leu417, Ile380->Leu422, Met338->Met430, Ile389->Ile431, Ala392->Gly434, Ala395->Glu437

4. Conclusions

We have conducted an extensive analysis of the HPs of Listeria monocytogenes. Following genome analysis of pathogens, we observed numerous proteins whose functions remain uncharacterized. Comparative genomics indicates that these proteins are present across organisms but have not yet been functionally characterized to date. The first part of this study comprises proteins for which biochemical activity can be predicted with reasonable confidence using genomic sequence-based approaches. We listed the proposed functions of most HPs using sequence analysis tools. The second part includes structure determination and the development of a structure-function relationship. We have successfully determined the structure-function relationship for 13 HPs from Listeria monocytogenes. This structural genomics project enables us to better understanding of pathogenesis of this microorganism. We have finally achieved our goal of assigning structure-based biochemical functions to most hypothetical proteins from Listeria monocytogenes. However, further biochemical work is needed to assign functions with accurate precision.

5. Conflict of Interest:

There is no conflict of interest to declare.

6. Data Availability Statement:

The data supporting this study are provided in this article.

7. Funding:

None

8. Acknowledgements:

None

9. Supplementary materials:

None

10. Author Contributions:

S.K. conceptualized and designed the study. S.K. performed data curation, and all reported analysis. S.K. conducted formal analysis and interpreted the results. S.K. prepared figures and tables and wrote the original draft of the manuscript. The author reviewed and approved the final version of the manuscript.

References

  1. K. H. Chin, Y. D Tsai, N. L. Chan, K. F. Huang, A. H. J. Wang, S. H Chou. The crystal structure of XC1258 from Xanthomonas campestris: a putative procaryotic Nit protein with an arsenic adduct in the active site. Proteins. 69. 665-671. 2007. https://doi.org/10.1002/PROT.21501

  2. G. Fibriansah, A. T. Kovacs, T. J. Pool, M. Boonstra, O. P. Kuipers, A. M. W. H Thunnissen. Crystal Structures of Two Transcriptional Regulators from Bacillus cereus Define the Conserved Structural Features of a PadR Subfamily. PLOS ONE. 7. e48015. 2012. https://doi.org/10.1371/JOURNAL.PONE.0048015

  3. M. Y. Galperin, E. V Koonin. "Conserved hypothetical" proteins: prioritization of targets for experimental study. Nucleic Acids Research. 32. 5452-5463. 2004. https://doi.org/10.1093/NAR/GKH885

  4. E. Z. Gebremedhin, G. Hirpa, B. M. Borana, E. J. Sarba, L. M. Marami, K. A. Kelbesa, N. D. Tadese, H. A Ambecha. Listeria Species Occurrence and Associated Factors and Antibiogram of Listeria monocytogenes in Beef at Abattoirs, Butchers, and Restaurants in Ambo and Holeta in Ethiopia. Infection and Drug Resistance. 14. 1493-1504. 2021. https://doi.org/10.2147/IDR.S304871

  5. J. Gury, L. Barthelmebs, N. P. Tran, C. Divies, J. F Cavin. Cloning, Deletion, and Characterization of PadR, the Transcriptional Repressor of the Phenolic Acid Decarboxylase-Encoding padA Gene of Lactobacillus plantarum. Applied and Environmental Microbiology. 70. 2146-2153. 2004. https://doi.org/10.1128/AEM.70.4.2146-2153.2004

  6. Y. Ishigaki, X. Li, G. Serin, L. E Maquat. Evidence for a pioneer round of mRNA translation: mRNAs subject to nonsense-mediated decay in mammalian cells are bound by CBP80 and CBP20. Cell. 106. 607-617. 2001. https://doi.org/10.1016/S0092-8674(01)00475-5

  7. J. Liu, N. Oganesyan, D. H. Shin, J. Jancarik, H. Yokota, R. Kim, S. H Kim. Structural characterization of an iron-sulfur cluster assembly protein IscU in a zinc-bound form. Proteins. 59. 875-881. 2005. https://doi.org/10.1002/PROT.20421

  8. G. Lubec, L. Afjehi-Sadat, J. W. Yang, J. P. P John. Searching for hypothetical proteins: Theory and practice based upon original data and literature. Progress in Neurobiology. 77. 90-127. 2005. https://doi.org/10.1016/j.pneurobio.2005.10.001

  9. J. A. Newman, L. Hewitt, C. Rodrigues, A. Solovyova, C. R. Harwood, R. J Lewis. Unusual, dual endo- and exonuclease activity in the degradosome explained by crystal structure analysis of RNase J1. Structure. 19. 1241-1251. 2011. https://doi.org/10.1016/j.str.2011.06.017

  10. M. Nuesch-Inderbinen, G. V. Bloemberg, A. Muller, M. J. A. Stevens, N. Cernela, B. Kolloffel, R Stephan. Listeriosis Caused by Persistence of Listeria monocytogenes Serotype 4b Sequence Type 6 in Cheese Production Environment. Emerging Infectious Diseases. 27. 284. 2021. https://doi.org/10.3201/EID2701.203266

  11. S. Sivashankari, P Shanmughavel. Functional annotation of hypothetical proteins - A review. Bioinformation. 1. 335-338. 2006. https://doi.org/10.6026/97320630001335

  12. I. I. Vorontsov, G. Minasov, O. Kiryukhina, J. S. Brunzelle, L. Shuvalova, W. F Anderson. Characterization of the deoxynucleotide triphosphate triphosphohydrolase (dNTPase) activity of the EF1143 protein from Enterococcus faecalis and crystal structure of the activator-substrate complex. Journal of Biological Chemistry. 286. 33158-33166. 2011. https://doi.org/10.1074/jbc.M111.250456

  13. K. Wada, N. Sumi, R. Nagai, K. Iwasaki, T. Sato, K. Suzuki, Y. Hasegawa, S. Kitaoka, Y. Minami, F. W. Outten, Y. Takahashi, K Fukuyama. Molecular Dynamism of Fe-S Cluster Biosynthesis Implicated by the Structure of the SufC2-SufD2 Complex. Journal of Molecular Biology. 387. 245-258. 2009. https://doi.org/10.1016/j.jmb.2009.01.054

Genome-wide annotation and structural modeling of hypothetical proteins in Listeria monocytogenes | EditoryPress