Advanced Search
Last updated date: Dec 16, 2024 Views: 385 Forks: 0
Rumley, J.D., Kim, J.H., and Hobert, O. (2025). Protocol to identify transcription factor target genes using TargetOrtho2. STAR Protocols 6, 103680. https://doi.org/10.1016/j.xpro.2025.103680.
https://www.sciencedirect.com/science/article/pii/S2666166725000863?via%3Dihub
https://pubmed.ncbi.nlm.nih.gov/40056408/
This protocol describes the local installation and use of the transcription factor target identification program TargetOrtho2 on MacOS systems. Based on an in silico phylogenetic footprinting approach, TargetOrtho2 uses transcription factor binding site information to predict phylogenetically conserved transcription factor targets and has been validated for its usefulness in nematodes including C. elegans. To properly function, TargetOrtho2 requires installing a suite of programs. Here we provide basic installation instructions for all required programs, as well as instructions for running TargetOrtho2.
We describe here a protocol for the local installation and use of TargetOrtho2 on MacOS systems. TargetOrtho2 is a program that uses transcription factor binding site information to scan whole genomes for the occurrence of such sites, associating them with the upstream regions, introns, exons, and downstream regions of genes. Putative target genes are ranked by their likelihood to be true targets based on several transcription factor binding motif features, including predicted binding affinity but, most importantly, phylogenetic conservation of motifs in the upstream regions and introns of orthologous genes (“phylogenetic footprinting”)(Figure 1) 1,2 The program, as published, searches the genomes of up to eight nematode species for phylogenetically conserved transcription factor binding sites, and is adaptable to search other well-annotated genomes, such as those of various Drosophila species. Several C. elegans genetics studies have used or adapted TargetOrtho2 and its predecessor TargetOrtho since its initial publication.3–11 However, since the Galaxy webtool on which previous TargetOrtho versions were running is no longer available, local installations of TargetOrtho2 are necessary. Here we provide easy to follow instructions to install and use TargetOrtho2 locally.
Timing: 1 h
TargetOrtho2 requires the operating system macOS X version 10.11.6 (El Capitan) or later, Xcode command line tools, the MEME Suite version 4.12.0 or later12, BEDOPS version 2.3.30 or later13, and bedtools version 2.27.1 or later.14 TargetOrtho2 requires Python 2.7, including the modules sklearn and pandas. TargetOrtho2 does not function with Python 3. A flowchart of TargetOrtho2 usage is provided in Figure 2.
xcode-select --install
Note: Make sure this installation is complete before attempting any installations using MacPorts
2. To install the final version of Python 2.7, go to the Python release downloads site (https://www.python.org/downloads/release/python-2718/). Download the macOS 64-bit installer and follow the installation instructions.
Note: After installation of Python 2.7 is complete, check which version of Python is set as default by entering the following command:
python --version
If the default Python version is a version of Python 3, check which version of Python 2 is installed by entering the following command:
python2 --version
The Python 2 version should be a version of Python 2.7. If the default Python version is a version of Python 3, follow the notes that begin with the phrase “If Python 3 is set as the default Python version.”
3. To install the sklearn module, enter the following command in the Terminal, and enter the password for your user account:
sudo pip install scikit-learn==0.20.4
Note: If Python 3 is set as the default Python version, enter the following command:
sudo pip2.7 install scikit-learn==0.20.4
4. To install the pandas module, enter the following command in the Terminal, and enter the password for your user account:
sudo pip install pandas
This should install pandas version 0.24.2
Note: If Python 3 is set as the default Python version, enter the following command:
sudo pip2.7 install pandas==0.24.2
5. To install the MEME Suite, go to the MEME Suite download site (https://meme-suite.org/meme/doc/download.html) and download the latest version of the MEME Suite. Decompress the downloaded compressed file. Install the MEME Suite on your computer by following the instructions on the MEME Suite installation site (https://meme-suite.org/meme/doc/install.html?man_type=web).
Note: We recommend performing the Quick Install using MacPorts. To install MacPorts, go to the MacPorts installation site and follow the instructions for installing MacPorts on your computer’s operating system (https://www.macports.org/install.php).
CRITICAL: The MEME Suite program fimo must be made executable by copying it from the directory /opt/local/bin to the directory /usr/local/bin.15 To do so, enter the following command in the Terminal from the directory /opt/local/bin, and enter the password for your user account:
sudo cp fimo /usr/local/bin
6. To install BEDOPS, go to the BEDOPS download site and download the installer package for OS X (https://bedops.readthedocs.io/en/latest/).
7. To install bedtools using MacPorts, enter the following command in the Terminal, and enter the password for your user account:
sudo port install bedtools
| REAGENT or RESOURCE | SOURCE | IDENTIFIER |
| Antibodies | ||
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| Bacterial and virus strains | ||
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| Biological samples | ||
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| Chemicals, peptides, and recombinant proteins | ||
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| Critical commercial assays | ||
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| Deposited data | ||
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| Experimental models: Cell lines | ||
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| Experimental models: Organisms/strains | ||
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| Oligonucleotides | ||
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| Recombinant DNA | ||
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| Software and algorithms | ||
| macOS X version 10.11.6 (El Capitan) or later | Apple | https://www.apple.com/app-store/ |
| Xcode command line tools | Xcode | N/A |
| MEME Suite version 4.12.0 or later | Bailey et al.12 | https://meme-suite.org/meme/doc/download.html
https://meme-suite.org/meme/doc/install.html?man_type=web |
| BEDOPS version 2.3.30 or later | Neph et al.13 | https://bedops.readthedocs.io/en/latest/ |
| bedtools version 2.27.1 or later | Quinlan et al.14 | https://bedtools.readthedocs.io/en/latest/content/installation.html |
| Python 2.7 | Python | https://www.python.org/downloads/release/python-2718/ |
| sklearn | Python | https://pypi.org/project/scikit-learn/0.20.4/ |
| pandas | Python | https://pandas.pydata.org/pandas-docs/version/0.24/install.html |
| TargetOrtho2 | Glenwinkel et al.1 | https://github.com/loriglenwinkel/TargetOrtho2.0 |
| execute_copy_terminal.py | This manuscript | N/A |
| TargetOrtho_motif_match_motif_search_terminal.R | This manuscript | N/A |
| TargetOrthoFIMO_motif_search_terminal.R | This manuscript | N/A |
| Python 3 | Python | https://www.python.org/downloads/ |
| R | R-project | https://cran.r-project.org/bin/macosx/ |
| Other | ||
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
| N/A | N/A | N/A |
Installation of TargetOrtho2
Timing: 10 min
This step instructs how to install TargetOrtho2 on a local MacOS computer.
python2 setup.command
6. Slightly modify the targetortho.py script.
a. Open targetortho.py in a text editor (e.g. TextEdit or BBEdit) or an integrated development environment (IDE; e.g. PyCharm, IDLE, or Xcode).
b. Use the Find function to locate all three instances of ‘# motif_id’ and change them to ‘motif_id’.
Timing: variable
This step instructs how to run TargetOrtho2 to predict transcription factor binding sites and transcription factor target genes in the C. elegans or P. pacificus genomes.
7. To run TargetOrtho2, in the Terminal, navigate to the directory TargetOrtho2.0. Run the script targetortho.py by entering a command such as the following in the Terminal:
python targetortho.py -f <input_file>
a. If Python 3 is set as the default Python version, enter the command as below:
python2 targetortho.py -f <input_file>
b. Input files for TargetOrtho2 are Position Specific Scoring Matrices (PSSMs) in MEME format (Figure 3). TargetOrtho2 accepts either log-odds matrices or letter-probability matrices. Troubleshooting 1
i. If your PSSM is in Cis-BP format (or another format), it must be converted to MEME format. To convert all PSSMs in a specified directory from Cis-BP format to MEME format, use the script execute_copy_terminal.py included in the supplemental materials (This script must be run with python3). To run execute_copy_terminal.py, enter the following command in the Terminal:
python3 execute_copy_terminal.py <input_folder> <output_folder>
ii. To convert individual PSSMs from a raw matrix to MEME format, use the script matrix2meme included in the MEME suite. Copy matrix2meme to the directory /usr/local/bin to make it executable. Enter the following command in the Terminal, and enter the password for your account:
sudo cp /Users/<username>/<meme_version>/scripts/matrix2meme /usr/local/bin
To run matrix2meme, enter the following command in the Terminal:
matrix2meme < <input_file>>><output_file>
c. Entering the command at the beginning of step 7 will run targetortho.py using the indicated input file, the default maximum p-value of 0.0001, and the default reference species of C. elegans. Troubleshooting 2
d. Other options to run the targetortho.py script are as follows:

i. By default, the output directory is titled <JobID>_TargetOrtho2.0_Results.
ii. The species file is a text file with a list of nematode species the genomes of which are searched for transcription factor binding sites and used to rate the likelihood of transcription factor target genes. By default, these species are Caenorhabditis elegans, Caenorhabditis briggsae, Caenorhabditis brenneri, Caenorhabditis remanei, Caenorhabditis japonica, Pristionchus pacificus, Pristionchus exspectatus, and Ascaris lumbricoides. Including a species file allows only a subset of these species’ genomes to be used in the analysis. In the species file, species names must be abbreviated as follows: c_eleg, c_brig, c_bren, c_rema, c_japo, p_paci, p_exsp, a_lumb.
All output files are in the output directory. For the purpose of identifying probable target genes of a transcription factor of interest, the most relevant output file is <JobID>_TargetOrtho2_ranked_genes_summary.csv. This file ranks genes based on Classifier Label Probabilities (class_prob). Classifier label probabilities of 0 to 1 indicate that a gene is increasingly likely to be a true target of the transcription factor of interest, and classifier label probabilities of 0 to -1 indicate that a gene is increasingly unlikely to be a true target.
In order to identify specific predicted transcription factor binding sites within target genes, use the files in the directory motif_match_data_per_species. The files in this directory are spreadsheets that list the predicted binding sites for the transcription factor of interest throughout the genomes of each of the searched species. The predicted binding sites are associated with the genes in the loci of which they are located. The sites are indicated as being upstream of the coding region, in an intron, in an exon (including untranslated regions), or downstream of the coding region, and its genomic coordinates are given. The quality of each site is indicated by a PSSM score and a p-value, based on how well each site conforms to the consensus sequence of the binding site.
An alternative method to identify specific predicted transcription factor binding sites is to use the files in the directory fimo_out. These files contain a list of all predicted binding sites for the transcription factor of interest associated with their genomic coordinates, but not the gene with which they are associated. These files also indicate the quality of each site with a PSSM score and a p-value that are based on how well each site conforms to the consensus sequence of the binding site. Troubleshooting 3
Most output files can be opened using Microsoft Excel or another spreadsheet program. For large output files, however, which can be generated if the p-value threshold is set high (weakly stringent), one can filter for the sites of interest based on either the associated gene name or the genomic coordinates of interest using the R scripts TargetOrtho_motif_match_motif_search_terminal.R or TargetOrthoFIMO_motif_search_terminal.R, respectively. These scripts are included in the supplemental materials. To run these scripts, enter the following commands in the Terminal, respectively:
Rscript <file_path>/TargetOrtho_motif_match_motif_search_terminal.R <input_file> <reference_gene>
Rscript <file_path>/TargetOrthoFIMO_motif_search_terminal.R <input_file> <chromosome/contig/scaffold> <start_search_position> <stop_search_position>
TargetOrtho2 is designed to function on Linux systems as well as on MacOS systems. This protocol, however, only describes its installation and use on MacOS systems, because we were not able to install the Python 2.7 sklearn and pandas modules on a Linux system.
Some input PSSMs result in empty fimo output files. This is due to the p-value threshold being set too low (stringent) for the particular input motif. To resolve this error, set the p-value threshold higher.
If you receive an error message indicating that a necessary Python module, such as scipy, is not installed, install the indicated module. To install scipy, enter the following command in the Terminal:
sudo pip install scipy
If one experiences difficulty in installing the full TargetOrtho2 program, and one’s main interest in using the program is to predict specific transcription factor binding sites, rather than predicting probable target genes of a transcription factor of interest, one can simply run the fimo portion of the script and manually identify the genes associated with sites of interest based on their genomic coordinates. In order to use TargetOrtho2 in this way, one must comment out clear_error() in line 894 in targetortho.py, otherwise the output files will be deleted when the program encounters an error.
Lead contact
Further information and requests for resources and reagents should be directed to and will be fulfilled by the lead contact, Oliver Hobert (or38@columbia.edu).
Technical contact
Technical questions on executing this protocol should be directed to and will be answered by the technical contacts, Jonathan D. Rumley (jdr2203@columbia.edu) and Jee Hun Kim (jk4213@columbia.edu).
Materials availability
This study did not generate new unique reagents.
Data and code availability
The python script execute_copy_terminal.py and the R scripts TargetOrtho_motif_match_motif_search_terminal.R and TargetOrthoFIMO_motif_search_terminal.R are included in the supplementary materials.
We would like to thank Zhenying Tian for writing the first version of the R script TargetOrthoFIMO_motif_search_terminal.R and for further work in developing the R scripts. We would also like to thank Zhenying Tian and Daniel M. Merritt for their help in troubleshooting the installation of TargetOrtho2. We would like to thank Marion Boeglin and Surojit Sural for allowing us to test the installation and use of TargetOrtho2 on their computers. This work was funded by the Howard Hughes Medical Institute (O.H.) and by a BRAIN Initiative NRSA F32 fellowship from the National Institute of Neurological Disorders and Stroke (F32MH136667; J.D.R). Jee Hun Kim was funded by grant R35GM131746 from the National Institute of General Medical Sciences (P.I. Iva Greenwald).
Conceptualization, O.H.; methodology, J.D.R, J.H.K, O.H.; investigation, J.D.R., J.H.K.; formal analysis, J.D.R., J.H.K.; writing – original draft, J.D.R.; writing – review and editing, J.H.K., O.H.; supervision, O.H.; funding acquisition, J.D.R., O.H.
The authors declare no competing interests.
During the preparation of this work the author Jee Hun Kim used ChatGPT4o in order to write the python script execute_copy_terminal.py. After using this tool/service, the authors reviewed the content as needed and take full responsibility for the content of the published article.
1. Glenwinkel, L., Taylor, S.R., Langebeck-Jensen, K., Pereira, L., Reilly, M.B., Basavaraju, M., Rafi, I., Yemini, E., Pocock, R., Sestan, N., et al. (2021). In silico analysis of the transcriptional regulatory logic of neuronal identity specification throughout the C. elegans nervous system. eLife 10, e64906. https://doi.org/10.7554/eLife.64906. PMID: 34165430.
2. Glenwinkel, L., Wu, D., Minevich, G., and Hobert, O. (2014). TargetOrtho: A Phylogenetic Footprinting Tool to Identify Transcription Factor Targets. Genetics 197, 61–76. https://doi.org/10.1534/genetics.113.160721.
3. Budirahardja, Y., Tan, P.Y., Doan, T., Weisdepp, P., and Zaidel-Bar, R. (2016). The AP-2 Transcription Factor APTF-2 Is Required for Neuroblast and Epidermal Morphogenesis in Caenorhabditis elegans Embryogenesis. PLOS Genetics 12, e1006048. https://doi.org/10.1371/journal.pgen.1006048.
4. Cornwell, A.B., Zhang, Y., Thondamal, M., Johnson, D.W., Thakar, J., and Samuelson, A.V. (2024). The C. elegans Myc-family of transcription factors coordinate a dynamic adaptive response to dietary restriction. GeroScience 46, 4827–4854. https://doi.org/10.1007/s11357-024-01197-x.
5. Masoudi, N., Tavazoie, S., Glenwinkel, L., Ryu, L., Kim, K., and Hobert, O. (2018). Unconventional function of an Achaete-Scute homolog as a terminal selector of nociceptive neuron identity. PLOS Biology 16, e2004979. https://doi.org/10.1371/journal.pbio.2004979.
6. Weinberg, P., Berkseth, M., Zarkower, D., and Hobert, O. (2018). Sexually Dimorphic unc-6/Netrin Expression Controls Sex-Specific Maintenance of Synaptic Connectivity. Current Biology 28, 623-629.e3. https://doi.org/10.1016/j.cub.2018.01.002.
7. Berghoff, E.G., Glenwinkel, L., Bhattacharya, A., Sun, H., Varol, E., Mohammadi, N., Antone, A., Feng, Y., Nguyen, K., Cook, S.J., et al. (2021). The Prop1-like homeobox gene unc-42 specifies the identity of synaptically connected neurons. eLife 10, e64903. https://doi.org/10.7554/eLife.64903. PMID: 34165428; PMCID: PMC8225392.
8. Kratsios, P., Pinan-Lucarré, B., Kerk, S.Y., Weinreb, A., Bessereau, J.-L., and Hobert, O. (2015). Transcriptional Coordination of Synaptogenesis and Neurotransmitter Signaling. Current Biology 25, 1282–1295. https://doi.org/10.1016/j.cub.2015.03.028.
9. Vidal, B., Gulez, B., Cao, W.X., Leyva-Díaz, E., Reilly, M.B., Tekieli, T., and Hobert, O. (2022). The enteric nervous system of the C. elegans pharynx is specified by the Sine oculis-like homeobox gene ceh-34. eLife 11, e76003. https://doi.org/10.7554/eLife.76003.
10. Reilly, M.B., Tekieli, T., Cros, C., Aguilar, G.R., Lao, J., Toker, I.A., Vidal, B., Leyva-Díaz, E., Bhattacharya, A., Cook, S.J., et al. (2022). Widespread employment of conserved C. elegans homeobox genes in neuronal identity specification. PLOS Genetics 18, e1010372. https://doi.org/10.1371/journal.pgen.1010372.
11. Sural, S., and Hobert, O. (2021). Nematode nuclear receptors as integrators of sensory information. Current Biology 31, 4361-4366.e2. https://doi.org/10.1016/j.cub.2021.07.019.
12. Bailey, T.L., Boden, M., Buske, F.A., Frith, M., Grant, C.E., Clementi, L., Ren, J., Li, W.W., and Noble, W.S. (2009). MEME Suite: tools for motif discovery and searching. Nucleic Acids Research 37, W202–W208. https://doi.org/10.1093/nar/gkp335.
13. Neph, S., Kuehn, M.S., Reynolds, A.P., Haugen, E., Thurman, R.E., Johnson, A.K., Rynes, E., Maurano, M.T., Vierstra, J., Thomas, S., et al. (2012). BEDOPS: high-performance genomic feature operations. Bioinformatics 28, 1919–1920. https://doi.org/10.1093/bioinformatics/bts277.
14. Quinlan, A.R., and Hall, I.M. (2010). BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26, 841–842. https://doi.org/10.1093/bioinformatics/btq033.
15. Grant, C.E., Bailey, T.L., and Noble, W.S. (2011). FIMO: scanning for occurrences of a given motif. Bioinformatics 27, 1017–1018. https://doi.org/10.1093/bioinformatics/btr064.
16. Nitta, K.R., Jolma, A., Yin, Y., Morgunova, E., Kivioja, T., Akhtar, J., Hens, K., Toivonen, J., Deplancke, B., Furlong, E.E.M., et al. (2015). Conservation of transcription factor binding specificities across 600 million years of bilateria evolution. eLife 4, e04837. https://doi.org/10.7554/eLife.04837.
Figure 1. Principle of TargetOrtho2. Shown here is an example output of TargetOrtho2 for the COE motif bound by UNC-3. The main output file is the target gene prediction summary file (ranked_genes_summary.csv), which shows a ranked list of predicted target genes in the reference genome. Genes are ranked by class probability, which is calculated by a Gaussian process classifier based on the motif features indicated in the table for sites in upstream regions and introns for each gene, as well as the alignment-independent conservation of motifs in upstream regions and introns in orthologous genes. Class probability ranges from –1 to 1, with higher values indicating greater likelihood to be a target of the transcription factor of interest. TargetOrtho2 also outputs motif match data files for each species with a list of detected motifs assigned to protein-coding genes and the region of each locus in which the site is detected. Other information about each motif is given, as in the table. The bottom portion of this panel illustrates the motif features used to predict transcription factor target genes. Shown is a schematized version of the results for COE sites in the zig-1 locus. TargetOrtho2 does not produce a graphical output.
Figure 2. Flowchart showing the pipeline of TargetOrtho2 usage. The input is one to five MEME format PSSMs in a text file. The program FIMO from the MEME suite searches all eight nematode genomes for motifs matching the first PSSM, and produces a genome hits table, which is output to the fimo_out folder. If more than one PSSM was in the input file, FIMO searches the genomes for that motif and adds these to the genome hits tables. BEDOPS software matches these genome hits to annotated protein coding genes and upstream regions, introns, exons, and downstream regions for each locus. Orthologous genes from the motif-gene association tables are matched, and the results are output in the motif_match_per_species folder. A Gaussian process classifier is used to rank predicted target genes based on motif features and their alignment-independent conservation in upstream regions and introns of orthologous genes. The results are outputted in the file ranked_genes_summary.csv.
Figure 3. Examples of MEME format PSSMs. These PSSMs should be .txt files. The sections of the PSSMs are as follows: i) The MEME version from which the PSSM file was produced. This does not have to be accurate for the proper functioning of TargetOrtho2. ii) The alphabet used for the PSSM. For DNA this is “ACGT”. iii) The strands in which to search for motifs. This should be “+ -” for searching both + and - DNA strands. iv) The background frequency of each DNA base to expect in the reference genome. v) The name of the motif. vi) Information about the PSSM, most importantly the type of matrix (log-odds or letter-probability), the number of classes of letters to expect (alength; 4 for DNA), and the number of letters to expect in the motif (w). vii) The PSSM with four columns (one for each DNA base in alphabetical order; i.e. ACGT) and the same number of rows as the number of bases in the motif, with each row representing successive positions in the motif. The values in each row are representations of the probabilities of each of the bases appearing at each position.1,16
Related files
All Figures.pdf
execute_copy_terminal.txt
TargetOrtho_motif_match_motif_search_terminal.txt
TargetOrthoFIMO_motif_search_terminal.txt Do you have any questions about this protocol?
Post your question to gather feedback from the community. We will also invite the authors of this article to respond.
Share
Bluesky
X
Copy link