Welcoming 2026 – independent research retrospective

2025 is gone, bringing with it another opportunity to share some highlights from a warehouse worker’s swing at studying our living history a sheaf of genome graphs and shelves full of protein-forms at a time.

Last year was both series of great adventures and turnabouts recorded in incomplete pieces of notes and impressions strewn out through the months. I think this is a good time to finally pastiche together those scraps into a travelogue of the year – maybe there’s another independent researcher out there someday who might find these experiences interesting.

Here are the two major highlights of the year.


ASM Microbe 2025: June 19-23

Last thing I remember before waking up was gently rocking heave-ho of the massive diesel train slowly pulling itself out from among the labyrinthine tunnels beneath white marble slabs of the new Moynihan hall, all grinding sound of metals and chains echoing through the dark.

Storm swept summer Kansas plains rolled by in a vivid haze of great horizontal lines outside the window – gold, brown and green under permanent grey cast of clouds trailing a sizeable tornado across the state. I didn’t expect what should have been a monotonous scene to be so overpowering, showering us with scent of the storm and grass tamed by the train’s HVAC. Unexpected transition of the train into a tornado chaser, however, meant our route into Colorado was blocked with precariously swinging power lines and trees strewn about, leading to ample delay (of about 24 hours) and chances to enjoy gamut of conversations among misfits in the cabins consisting of school teachers, truckers, civil engineers and a big group of Amish travelers that later broke out into songs over a bit of accordion. It was a welcome reprieve from gawking at shifting ecotopes and struggling to make my way through the Notebook of Malte Laurids Brigge, which was somehow proving to be a denser read than the Magic Mountain at less than half the volume.

Unexpectedly, these conversations proved to be a high point of the journey (I also had a lot of fun cactus-sighting during the Arizona leg of the trip – turns out we were just a bit too north for a good Saguaro spotting). There’s nothing quite like hearing about the Colorado River Compact from two retired & retiring civil engineering veterans from Nevada and California sitting across a table by pure happenstance, excitedly discussing current state of the Hoover dam reservoir like some erstwhile actors from the set of Chinatown. The delve into the unique bit of Americana was punctuated by moments of occasional silence from the gentleman from California, turning his coffee mug while staring into it with creased eyebrows – I believe he was en route to meet his son from Mexico.

These random chats gave me much needed opportunity to share my research to series of fresh audience as well, held captive by Amtrak steak dinner and desserts. A more memorable impromptu presentation of my ASM 2026 poster was with a trucker couple from Illinois – proudly having traveled to the 49 states (sans Hawaii) on their own wheels, they were taking it easy aboard a train for this journey. After a roller-coaster of threads ranging from their support for rounding up illegal immigrants, questions about latest covid conspiracy theories and difference between Archaea and Bacteria, I gave a somewhat condensed version of the poster-side chat with some neat finger drawings.

This one's more of a small demo of the Archaeal sequencing and analysis data I've been building up since last year - covering the whole Halococcus genus! 

All out of pocket, lifting boxes and selling bagels out of an aluminum box working at my own converted warehouse lab. It's looking like a genome library for a whole evolutionary branch of Archaea is set to balloon from a warehouse lab contribution by the end of the year.

You see, last year an academic partner and I found that a species of Halococcus, H. dombrowskii, carries an rrn operon on both its chromosome and plasmid (an rrn operon is just an operational unit for expressing ribosome - it's a protein we use as a signpost to figure out how related one species is to another - for now). This is an expansion of the study to see if the chromosome-plasmid distribution repeats across the entire genus or if it's specific to ours. 

On top of that, I'm gathering additional data on methylation state of the new genomes while sequencing them (methylation's a chemical modification on the DNA itself, decides whether the thing actually does something or not), since it's supported out of the box with latest flowcells. These genome sequencers are about yea big - measured in inches and fits in a pocket. Darn expensive to run though. It looks like methylation pattern on the chromosomal rrn differs from that found on the plasmid, which I'm really excited about.   
MinION sequencer sketch by Garin, one of the series being prepared for a research zine under Binomica name. CC BY-NC-SA 4.0

It sounded like they were quite enthused by the aspect of a warehouse worker’s presentation at an academic conference (maybe more so than the study itself) – “go get’em tiger” they said, among other encouraging things.

Abstract:

Euryarchaeon Halococcus dombrowskii has an unusual genomic architecture with multiple rrn (ribosomal RNA) operons distributed across its chromosome and plasmids. Our previous study on differences between H. dombrowskii’s plasmid and chromosome rrn operon ITS (Internal Transcribed Spacer) sites suggested there could be further quantitative differences among these disparate rrn operon regions. As the rrn operon ITS regions in question occur close to a well-characterized tRNA site known to interact with archaeal RNase, we performed an exploratory methylation survey of the H. dombrowskii genome to check for genomic compartment-specific rrn operon methylation patterns and site differences.

Here, we present preliminary long-read based direct methylation calling data for H. dombrowskii ATCC BAA-364, describing a genome rich in N4-methylcytosine (4mC) followed by N6-methyladenine (6mA). Initial survey describes a minor methylation site difference between regions in and around chromosomal and plasmid rrn operons that will require further analysis. We also note a curious lack of N4-methyltransferase annotation for the H. dombrowskii genome, and offer a putative candidate gene for further study based on HMM profiling and sequence homology analysis.

We hope to utilize the data and tools we developed for this initial study on follow-up long-read sequencing and direct methylation calling analysis of other Halococcus spp. Additional data would allow us to perform vitally needed comparative genomics analysis, providing us with further fascinating insight into archaeal ribogenesis in genomes with multi-compartment rrn operon placement.

Poster available at:

https://doi.org/10.5281/zenodo.16369513

After a quick engine failure at the entrance of the Mojave desert and some jumping wire fences and buses, I finally made it to LA in time for last tram into the city – at the strike of midnight before beginning of the conference.

Conference experience is far too long and complex to go through here – I was certainly happy to see familiar faces, many of them even welcoming. Quite a bit of valuable input on how to shape the Halococcus data as well. The discussion on pangenomic approach is well reflected in a preprint I’m writing up at this moment, set for final QC and submission this Spring.

I even got to rerun the poster spiel with more than a few people. One postdoc responded with her mouth agape in horror when I got to the warehouse and bagel part – “Oh I’m sorry!”


Boston Bacterial Meeting 2025: June 09-10

This one was on a whim. My first complete genome sequencing project after buying the MinION (there were other incomplete ones before) was that of Deinococcus radiophilus ATCC 27603 undertaken with another amateur biologist friend in 2019.

In retrospect, the sequencing was the straightforward part. Considering we had to spend months to scrounge up our own HMW genome extraction method from scratch in the dark era of ONT R9 chemistry, this is saying something. Beloved Flye was still in beta stages, and state of polishing tools were nowhere near what we have now. For moderately complex genomes, complete de novo assembly of acceptable quality just wasn’t very achievable up until Trycycler and Homopolish came on the scene (the latter’s replaced by many maintained alternatives, and Trycycler is superseded by Autocycler in most scenarios now – tooling’s come a long way).

It bears repeating – anything worthwhile in independent and amateur research would not have been possible without open source software and accessible hardware.

Figuring out genomic compartment level completeness of an assembly without a reference was pretty tricky to figure out in the beginning – and I had to learn about alignment, tree building and comparative genomics methods quite quickly. And then a method to track presence of highly variable genes across different genomes – in my case a task I had to learn basics of HMM and HMMER suite for. Tearing apart and studying well documented pipelines like GToTree was essential for this part. I still recommend GToTree (https://doi.org/10.1093/bioinformatics/btz188) and its documentation as a standout example of what an open source research software documentation should be.

At the time, studying completeness of a Deinococcus radiophilus genome meant understanding how other Deinococcus genomes are structured, and that eventually led to an interesting observation. It turns out vast majority of species in the Deinococcota phylum maintains large (100kb+) plasmids driven by a specific plasmid replication initiator protein (RPA) belonging to a single protein family with shared motifs and domains, representing close to 90% of the known genomes in the group (the percentage’s been steadily climbing as more genomes within the phylum are sequenced). So – for whatever the reason, the phylum Deinococcota represents the highest concentration of RPA proteins and RPA protein driven plasmids among 19 bacterial phyla surveyed from NCBI collection, spanning about 11k QC’d and deduplicated genomes (these numbers are significantly higher now – 11k genomes are from 2019 when I was just starting to learn about API access and other bioinformatics things).

The RPA proteins are always a component of an operon next to parA and/or parB genes, with recognizable palindromic sequence motif following the operon – so chances of the HMM screening results being entirely false positives is unlikely.

While further study of RPA distribution and evolution in nature was held off back in 2019 due to life circumstances, I thought it was about time I do a proper study of RPA evolution within the Deinococcota phylum. I wanted to further pursue my initial idea that unusually high concentration of RPA among Deinococcota is because they’re distributed through inheritance from an ancient common ancestor.

And what better way to present this work-in-progress than an early career/student researcher facing conference like Boston Bacterial Meeting?

Abstract:

Replication initiator protein A (RPA) under family PF10134 are often found in bacterial genomes as the main component for specific types of single strand DNA binding plasmid replication initiator sites, alongside accessory partitioning proteins A and B.

Screening 11082 reference bacterial genomes representing 19 phyla for RPA plasmid replication initiator protein (PF10134) and accessory partitioning genes show drastically uneven distribution of RPA plasmid operon in nature, with phylum Deinococcota hosting an exceptionally high concentration of the mechanism at 89% of member microbes representing extremes of habitat heterogeneity.

Multiple alignment and inferred tree experiments using Deinococcota RPA proteins show patterns mirroring rooted whole genome single copy gene (SCG) phylogenetic trees of their hosts. Current data suggests wide distribution of RPA based plasmid replication system exclusive among members of Deinococcota is based primarily on inheritance rather than introgression. Further study into evolution of RPA plasmids in Deinococcota could provide us with an unusual model system for studying plasmid evolution at scale.

Poster available at:

https://doi.org/10.5281/zenodo.16369073

The folks at BBM also offers scholarship to offset the registration costs – and fortunately thought my application was interesting enough for attendance. Much thanks to the nod, I honestly would not have been able to attend this one without the waiver.

The meeting was a fruitful one – multiple attendees suggested that I focus on analyzing the nature of conserved domains and motifs within the Deinococcota RPA proteins before proceeding further (I was originally considering expanding the screening outside the phylum to see if there’s an ‘ingress point’ protein out there somewhere). Some new tools and databases were suggested as well, something I’ll address in a separate note addressing this study.

What surprised me about the meeting was an atmosphere of universal focus on trimming down the research question until an observation is a wetlab problem. This Deinococcus study, for example, would not be considered fully featured (publication worthy) unless I started doing transformation characterization of the RPA protein and some model plasmid carrying it. I mentioned that I’m not sure if the approach is within the purview of a study on potential RPA common ancestor for a phylum level group – but was left with an impression that a strictly tree building study would be lacking.

I think that’s a fully justifiable view – but how would one design a wetlab counterpart to a study like this? Understanding the RPA on a more mechanical level should provide the answer, and is currently my focus on trimming and redirecting this study.

A memorable session at the conference was an elective titled ‘Science for the Public Good.’ I wasn’t quite sure what to expect wading my way through the posters and crowds, only to find the classroom packed with audience with an almost palpable sense of curiosity and excitement in the air. As an outsider the choice of most hotly debated topic during the session remains curious to me (graduate student unionization), especially within the context of public good. Even within the scope of graduate student well-being, isn’t it natural to also consider separating the identity of the scientist from their institutional affiliations in pursuit of agenda separate from those of the funding sources and funding administrators?

I got the sense that contemporary academia in the US largely considers itself under siege from what they consider outsider elements, such as the US government. However, is this really an accurate perspective when we consider the model of post-WWII scientific research built up under leadership of people like Vannevar Bush that we still operate under – where science IS the government/university complex? Again, it was fascinating to observe a cultural gap like this firsthand – this is something I’m aiming to address fully sometime in 2026.

Another notable experience was running into some people from SeqCoast sequencing company. Inexplicably, multiple greetings and questions to one of their senior personnel were met with pointed silence and staring off into the thin air. It’s been a while since I’ve seen someone try to ignore someone else with such intense effort. Although one of them came up later for a chat, past years of identifying myself as an independent researcher taught me to give a wide berth with organizations like this. What an odd experience, one would think people would be more talkative at academic conferences.


2025 in numbers

Number of genomes sequenced12
Number of genomes submitted6
Number of preprints written/in progress4
Number of collaborations3
Number of coauthorships2*

This post/note was originally a lot longer – addressing some of the extremely disheartening, personally damaging experiences with academia this year as well as my thoughts on requirements & attitude for an independent researcher that began to emerge throughout the year. However, those thoughts deserve a more complete, separate coverage.

It would be best to close out the last day of 2025 looking back on the more hopeful things, of solid discoveries of previously unknown things in this great, beautiful universe.

I genuinely believe there is a room for the identity of the scientist separate from other social institutions, and that everyone with the passion and drive can try their hand at scientific discovery in the same way being a poet or a novelist is not tied to our social backgrounds and affiliations. It’s just the question of building the way.

Happy 2026 everyone, the year of the Fire Horse!

Setting up GNU Guix package manager

It’s been a hot minute since last update on this lab note page. Things since last write up had been quite hectic, both in good and bad sense. I hopped on a train across continental US to attend (and present: https://doi.org/10.5281/zenodo.16369513) at ASM 2025, went through significant bit of abuse as a sequencing and analysis technician for sequencing core of the New York Medical College, and finally presented a smaller bit of an ongoing plasmid replication gene evolution study at Boston Bacterial Meeting 2025 (https://doi.org/10.5281/zenodo.16369073)

It was an eye opening experience for sure – reeling from the experience kept me away from writing here as well, but that is something I’ll have to address in a separate write up down the road.

I’ll just say this one thing – when working in proximity of people in the academia and its institutions, it would be very prudent for any aspirational amateur or independent researchers to be always aware of their surroundings and leave thorough paper trail.


GNU Guix (https://guix.gnu.org/) is an aspirational topic of study for me. I first became aware of it around 2020 while looking for workflow language candidates for a project involving screening of plasmid replication initiation genes across SRA raw reads data (that later went on to become the topic of BBM 2025 poster linked above). From the beginning it was apparent that an effective, reproducible biology analysis workflow would require a minimal reproducible encapsulation of the working environment, not just strung together series of commands.

As of the time of this writing (2025 October) the question is addressed with widespread usage of conda environment and venv/pip combination (or through uv, which combines some aspects of the two in complementary manner). However, the virtual environment methods have their detractors, enough for many people to consider fully functional, atomic-rollback capable, immutable package management systems such as GNU Guix and Nix as the way to go for the future. From what I’ve seen there’s an active interest in HPC community in this type of approach.

My first foray into GNU Guix was problematic at best, owing in large part to my inexperience with linux systems in general. After all, by that point I’ve only been a full-time linux user and somewhat-proficient scripter for about a year or so. Rather than being able to translate my messy bash workflow script into guile scheme on top of Guix, I ran into constant issues just getting Guix to sit on top of the host distro properly. Combined with both time requirement and difficulty in iterating through multiple installation of Guix (best way to practice with these sorts of systems, in my experience, involves being able to remove the whole thing cleanly and then re-iterate with different patterns from ground up), I had to put the experiment on hold indefinitely.


Fast-forward to 2025, I’ve been doing this long enough to know I’ll be studying biology for the rest of my life, and my projects are getting larger and more complex. Getting broader breadth and depth of experience in practices like GNU Guix is no longer just a passing curiosity, but a fundamental requirement of a lifelong study.

Conda/venv/pip (and now uv, as mentioned before) are my everyday tools and I’ve been delivering some work in docker containers already – but is this all there is? Maybe it’s finally time to go back, setup GNU Guix properly and try to understand it. Heck, maybe I’ll even pick up some emacs and guile scheme skills while I’m at it!

This is the first of my notes toward studying the Guix system, alongside learning guile scheme as a sort of elaborate systems scripting alternative.


I’m focusing on working with GNU Guix as a package manager that sits on top of a host distro.

Quick note – I tried out below processes on Ubuntu, Debian (via LMDE 7) and Fedora 34 beta, and it looks like SELinux distros (such as Fedora and openSUSE) are not as plug-and-play as the alternatives. The installation script should setup alternate policies during installation (as described here https://guix.gnu.org/manual/devel/en/html_node/SELinux-Support.html) but there was no end to read-write and permission related issues. It should be possible to setup Guix on Fedora, since there are obviously people using the package manager on the distro – I just didn’t have the time to figure it out when Guix works just fine on Ubuntu and LMDE 7 as far as I know.

The default GNU Guix installation script is very much improved from what I remember – notably it offers a very comprehensive uninstall option. I’ve deliberately uninstalled working GNU Guix setups and checked the system for any remaining files or directories, and the script really does its job allowing for retries without any worry that you might break your system (of course, I’d still recommend doing all this on an extra personal machine first if you have the option).

Lastly, this is a GNU Guix study note from someone new to the system! What’s written here isn’t an expert advice or guide by any means. When in doubt, #guix on mastodon seems to be very actively read by experienced members of the community.


Step 1 – let’s download the installation script, chmod it and install it as superuser.

wget https://guix.gnu.org/install.sh
chmod +x install.sh
sudo ./install.sh

Unless there’s a special reason, just answer yes to all the prompts. The script will ask us if we want to import the necessary GPG keys and whether we want precompiled binaries or recompile every package – we definitely want precompiled binaries.

Step 2 – we need to add GUIX_PROFILE entries in our .profile, located in our home directory. I guess it’s possible to use .bashrc as well, but these are paths we’ll want on per-login basis rather than per terminal session basis.

Add below paths to the .profile file.

GUIX_PROFILE="$HOME/.guix-profile"
. "$GUIX_PROFILE/etc/profile"
GUIX_PROFILE="$HOME/.config/guix/current"
. "$GUIX_PROFILE/etc/profile"

What’s happening is we have two GUIX_PROFILE, one for installed packages, and another one for Guix system itself. There’s the home directory .guix_profile directory for our applications, and .config/guix/current directory for the user facing Guix system. Both are required to have a working Guix installation.

After adding above lines, source the profile within the terminal session via below command, assuming we’re in home directory. This will apply above paths in the .profile for the current session.

source .profile

By now it’s a good idea to make sure the guix-daemon is running in the system. This is the GNU Guix component that will do all the writing and reading of new package and its dependencies into unique, hashed directories. Assuming systemd based distro, you can check it via:

sudo systemctl status guix-daemon

Step 3 – let’s install a demo package called ‘hello’, which prints hello world message to the terminal when executed. It’s a nifty little package to check if the setup so far is working properly or not.

Quick note – one unexpected issue I’ve sometimes run into (once on Debian, maybe also on Ubuntu? I’m not too sure), is that .guix-profile directory referred to in above .profile file isn’t generated automatically during the installation process. I had to try to install a package via Guix to have it generate the directory.

This could lead to some unexpected weirdness, especially on LMDE 7. Apparently .profile calling on file/directory that does not exist can result in a login screen loop after a system restart/logging out. Meaning the user would enter their password and try to login, only to return to the same login screen ad nauseam.

If you do run into a similar issue, simply press ctrl+alt+f2 to log in via text terminal, and delete/comment out the offending line in the .profile and either reboot or press ctrl+alt+f7

guix install hello

This should install the ‘hello’ package you can activate from the command line by typing hello, and create .guix_profile directory if one wasn’t present before.

Step 4 – let’s update the Guix system itself. You can do so via the pull command.

guix pull

This will take quite a bit of time, at least the first time (which is why we started with guix install hello first).

Step 5 – we need to install glibc-locales… And this part is where I find myself diverging from many other tutorials out there.

Quick note – materials like the excellent System Crafters tutorial (https://systemcrafters.net/craft-your-system-with-guix/installing-the-package-manager/) call for sudo guix installing glibc-locale and modifying .profile at the system root level. For whatever the reason I find that I need to install glibc-locale and update .profile twice, first as a sudo and second as a regular user in order to make the change detectable by other Guix installed packages, such as guile scheme.

First, make sure you have su access (root password needs to be set). And then carry out below steps:

sudo guix install glibc-locales
sudo nano /root/.profile

And add below lines to /root/.profile

 GUIX_LOCPATH=$HOME/.guix-profile/lib/locale

And then do the same for regular guix session:

guix install glibc-locales
nano .profile

Adding the same GUIX_LOCPATH to home directory .profile

 GUIX_LOCPATH=$HOME/.guix-profile/lib/locale

And source the .profile as described before. We might want to reboot our system once here just to be on the safe side.

Step 6 – installing baseline packages. If you made it this far and are not running into any error messages (usual suspects include issues with PATH, read/write access, and lack of .guix_profile directory), you’re good to go.

I’m using GNU Guix as an opportunity to learn emacs and guile scheme as well, so I’ll be setting them up with readline and colorization support.

guix install emacs
guix install guile
guix install guile-readline
guix install guile-colorized

The guile packages readline and colorized are modules, and needs to be loaded via a .guile config file in the home directory. Chances are we’ll need to create one ourselves, containing below lines.

(use-modules (ice-9 readline)
    (ice-9 colorized))

(activate-readline)
(activate-colorized)


We have a working GNU Guix package manager running on top of our distro. Next time I’ll cover adding free and non-free scientific computing package repositories, Guix shell and home concept, and setting up emacs for a more elaborate guile scheme coding sessions.

Looking into Pangean relics

While on the usual internet lam traveling down the endless piles of papers reporting on fascinating new mysteries from the microbial world, an MRA stood out and caught my attention.

Genome sequence of an extremely halophilic archaeon isolated from Permian Period halite, Salado Formation in New Mexico, USA: Halobacterium sp. strain NMX12-1

Soto L, DasSarma P, Anton BP, Vincze T, Verma I, Eralp B, Powers DW, Dozier BL, Roberts RJ, DasSarma S. Genome sequence of an extremely halophilic archaeon isolated from Permian Period halite, Salado Formation in New Mexico, USA: Halobacterium sp. strain NMX12-1. Microbiology Resource Announcements. 2024 Nov 12;13(11):e00778-24.

Yet another Haloarchaea isolated specifically from a Permian halite! Much like our Halococcus dombrowskii (https://doi.org/10.1101/2022.08.16.504008 ) which is also said to be isolated from a Permian era deposit of an underground salt mine (https://doi.org/10.1099/00207713-52-5-1807).

Having gone through the sequencing and comparative analysis process for the H. dombrowskii, I’ve always remained skeptical of the species being truly isolated and preserved from a Permian era ecosystem. An issue in the case of H. dombrowskii was that we had plenty of non-halite captured samples of Halococcus species out in the wild maintaining reasonable & expected genomic distance from the supposed living fossil of a captured ecosystem from 290~250 million years ago. Here’s a quick cladogram of a phylogenetic tree (NNJ) built from alignment of Halococcus species derived single copy genes set:

Halococcus SCG set cladogram (NNJ)

H. morrhuae is the closest living cousin to H. dombrowskii, and while provenance of the specific species is a little vague (I found records dating back to 1880’s description of contamination in fisheries mentioning H. litoralis, also related to H. schoop, both sharing coccoid features and brick-red coloration under saline conditions), H. morrhuae certainly wasn’t isolated from our contemporary ecosystem for the last 250 million years.

H. morrhuae DSM 1307 micrograph. Intellectual property rights:
Leibniz-Institut DSMZ-Deutsche Sammlung von Mikroorganismen und Zellkulturen GmbH

And yet – people do keep finding Haloarchaea specifically around Permian era salt deposits somehow. Is this simply because researchers tend to specifically sample from Permian era deposits, or has there been multiple control dig and isolation attempts, and we keep on finding viable cells only from Permian era layers? I’d be very interested in reading additional follow-ups from the authors of the recent MRA.

Focus on geological salt layers from Permian era does have a sort of romantic rationality to it, of course. End of the Permian era was the Permian-Triassic extinction event, eventually (in MYA scale) followed up by breaking up of Pangea. And it’s widely assumed the catastrophic rearrangement of continents and following sea level change created salt deposits that persisted across different continents of our modern era.

Looking at Haloarchaea as possible remnants of that very special period in the Earth’s history is certainly something I like to think about. Imagine emergence of previously unknown shades of pink and red pools across the ur-continent as the Earth itself gradually rearranged itself. The new type of microbes would both gradually spread and also be captured in salt crystals and buried deep beneath the earth, only to be released again time to time throughout history.


While the MRA paper stood out due to its supposed Permian origin, the whole reason why I’m reviewing Haloarchaeal papers is because I’m undertaking a somewhat large sequencing and analysis project.

I’ve made some very exciting observations while assembling and analyzing Halococcus dombrowskii, described in bits and pieces in our initial biorxiv preprint and follow up poster presented at ASM Microbe 2024 (I should really remember to write on that one too). And as it stands only way to confirm the observation is to re-sequence and create high quality, closed genomes of other members of the Halococcus genus!

Currently the humble warehouse-lab of yours truly has five samples of Halococcus species ready for sequencing and analysis. Our previous experience with Halococcus was hampered by its somewhat ridiculous doubling time of two weeks – but the new media formulation suggested by a collaborator has sped up that time frame drastically. Granted, I’m still not getting the sort of clearly defined single colonies, but I’ll take what I can get.

Here’s a general gist of what the new media’s made of (adapted version of Hv-Cas medium described here https://pmc.ncbi.nlm.nih.gov/articles/PMC8131023/) :

372-Cab

1L
KCl 4.2g
MgCl2*6H2O 18g
MgSO4*7H2O 21g
NaCl 195g

After cooling to ~70C, add 12 mL of 1 M Tris-HCl, pH 7.4
And add:

10x Casamino Acids Stock 100ml
1000x Vitamins Stock 1ml
CaCl2 1M 3ml

Add 50ml of Trace Elements stock 

10x Casamino Acid Stock: 
Casamino acids 50g/L

1000x Vitamins Stock:
Thiamine (1 g/L) 1g
Biotin (0.1 g/L) 100mg

Trace Elements Stock (1L, store in 50ml aliquots):
EDTA 5g
FeCl3 (4.9 mmol) 0.8 g
ZnCl2 (0.37 mmol) 0.05 g 
CuCl2 (0.074 mmol) 0.01 g
CoCl2 (0.077 mmol) 0.01 g
H3BO3 (0.16 mmol) 0.01 g
MnCl2 (12.7 mmol) 1.6 g
NiSO4 (0.065 mmol) 0.01 g
Na2MoO4.2H2O (0.041 mmol) 0.01 g

And a quick reference on the exact type of trace elements used for the mix (based on the original paper linked above):

5 g/l Na2EDTA•2H2O (Sigma E1644) = EDTA disodium salt, dihydrate

0.8 g/l FeCl3 (Sigma 157740) = iron chloride, anhydrous

0.05 g/l ZnCl2 (Sigma 793523) = anhydrous

0.01 g/l CuCl2 (Sigma 751944) = copper chloride , anhydrous

0.01 g/l CoCl2 (Sigma 232696) = cobalt chloride, anhydrous

0.01 g/l H3BO3 (Sigma B6768) = boric acid 

1.6 g/l MnCl2 (Sigma 328146) = manganese chloride, anhydrous

0.01 g/l NiSO4 (Sigma 656895) = nickle sulfate, anhydrous

0.01 g/l Na2MoO4•2H2O (Sigma M1003) = sodium molybdate dihydrate

I’m not entirely sure how much of the trace elements are actually necessary – the biggest change here is frankly swapping out of the yeast extract for casamino acids. I’m currently testing out some sea-water type trace element mixes to save on cost and labor, and will be documenting any interesting observations here.

Nanopore: archiving signals

Last year’s experience working with MacLea lab brought with it quite the varied ensemble of microbial sequencing side-projects, involving both long and short read platforms.

Eventually we ended up settling into a routine of:

  1. Finding taxonomically curious samples within the particular taxonomic domain (our primary topic was phage polyvalence – so addressing host genetics was part of the research project).
  2. Screen through 16S only data in taxonomic periphery of the target genome. I mostly focused on Silva but occasionally had to dig through NCBI manually. The 16S based data would be compared against existing taxonomic classification based on traditional microbiology/biochemistry methods.
  3. Confirm any anomalies or mismatches and queue target microbes up for deeper analysis/sequencing.

All of which are mostly automated now by yours truly, and will be turned into a pipeline and written up pretty soon(ish).

Increased pace of sequencing meant more data to share. Before I would simply upload raw fastq files and call it a day, but rapid improvements in long read signal processing in recent years has increasingly turned fastq into a momentary snapshot rather than future-facing raw data useful for other researchers.

My own experience looking at older R9 chemistry ONT data seems to confirm this view. Downstream analysis improvements resulting from simply re-basecalling the original raw signal data using newer models and pipelines are becoming hard to ignore, so much so that I’m currently preparing a major update to our own D. radiophilus genome based on newer model.


Researchers working with ONT platform and looking to preserve their data for posterity (and fruitful future re-analysis) have three options when it comes to working with raw signal data.

  1. Fast5 – original raw data format from ONT based on HDF5 like structure (the team that went on to create blow5/slow5 format wrote what might be the most comprehensive public documentation on fast5 – available here: https://hasindu2008.github.io/slow5specs/fast5_demystified.pdf). This format is being rapidly phased out, replaced by…
  2. Pod5 – new raw data format from ONT based on Apache arrow-columnal format. With Dorado being the basecaller of choice for ONT, we’ll be looking at more pod5 data in the future
  3. Blow5/Slow5 – a third party format introduced a few weeks before ONT unveiled their pod5 format. Possibly the simplest structure of the three resembling SAM/BAM file with header portion and TSV like data format

For raw data generated by latest crop of experiments, I tried uploading the fast5 output to NCBI SRA as I’ve done before. Alas, my experience had been mixed. Apparently something about fast5 format from latest ONT pipeline is a bit different now, no longer encoding a predicted sequence information within the raw data file itself (which makes sense – we’re using the raw to basecall sequence information separately).

Unfortunately, NCBI’s fast5 upload pipeline seems to look specifically for the encoded sequence information as part of its QC and will reject raw data files without it (which I assume applies to all newer, R10 type raw data unless specific options are toggled – are too few people uploading ONT raw data for this to be an issue, or am I missing something?). And despite being the obvious up and coming ‘official’ data format of ONT platform, pod5 is not officially supported by NCBI at this time (July 2024).

Every crisis can be an opportunity – I’d use this one to figure out where & how to upload our current and future ONT raw data, and decided to throw in a twist. I wanted to figure out how to upload and archive an alternate, third party specification of ONT raw data.

Note the archiving aspect – as far as simple uploading of data is concerned I’m sure we can all figure out how to get things on AWS or some private FTP (or even a torrent) and call it a day. Yet we don’t just want to make the data available – we need to make them available for posterity, otherwise they would not be able to fulfill their intended purpose – an aid for any future researcher, perhaps even decades in the future, to pursue their own study. It’s a difference between throwing something up in an attic versus sending it off to a museum.

And then there’s the third party data specification aspect. Our sequencing technologies are becoming ever more sophisticated and further removed from analog chemistry of our subjects. While some of the advances and its benefits border on science fiction in its rapid pace, we might be slowly approaching a period where some serious consideration on open and compatible standards for deciphering the machine-to-raw DNA portion of the pipeline could be necessary. ONT is a for profit company – how long do they last on average? Imagine we live in a world where if Simon & Schuster goes out of business, third of their publications turn into gibberish – or at least tricky to read.

Using, and thus, driving an impetus to maintaining an alternate specification of these increasingly important data will at the very least keep us studying and documenting them. It’s a net positive for scientific research and human knowledge at large, in my humble opinion.


A twitter/mastodon discussion on how to upload raw data from ONT platform on archive-quality data repository caught the attention of Psy_Fer_, or James Ferguson, one of the co-authors of the Slow5 paper. Soon Hasindu, first author of the aforementioned paper graciously reached out and suggested I use ENA (European Nucleotide Archive) for raw data upload and offered to walk me through the upload process.


First, we have to use the slow5tools to convert existing fast5 raw data into blow5. Slow5 is the larger, human readable version of the raw data. Blow5 is the compressed, smaller file of the same thing that’s more suited for archiving, since these files can get ludicrously large very quickly.

#first, initial blow5 files were generate from fast5 via:
slow5tools f2s barcode20/ -d blow5_barcode_20_p
slow5tools f2s barcode20/ -d blow5_barcode_20_f

#and I'll merge all files in pass and fail directories, respectively, so I'm not keeping track of couple of hundred files at once:
slow5tools merge blow5_barcode20_p/ -o blow5merge_barcode20_pass.blow5 
slow5tools merge blow5_barcode20_f/ -o blow5merge_barcode20_fail.blow5 

#while I'm at it, I'll also generate index files for each of the merged blow5 data
slow5tools index *.blow5 

#once all the steps are done, make sure all the files and directories are organized as you'd like to see in the final version, and create a tarball for ENA upload.
tar cvf - ont_r10_5kHz_dna_a_sp_21022_blow5/ | pigz -O - > ont_r10_5kHz_dna_a_sp_21022_blow5.tar.gz

#a difference I observed between NCBI and ENA - ENA asks for MD5 checksum during the file upload step.
md5sum ont_r10_5kHz_dna_a_sp_21022_blow5.tar.gz 

Once the data is ready, create an ENA account if you don’t already have one here: https://www.ebi.ac.uk/ena/browser/submit (for future reference, proper login URL once you have your account is at https://www.ebi.ac.uk/ena/submit/webin/login)

From here on the general workflow goes:

  1. Register a study
  2. Register a sample
  3. Submit reads/data

Once the study is registered and the data is ready in its final archived form (we’ll need its checksum) it’s time to download ENA’s metadata template and fill it out for the sample registration step. A caveat here is that we need to choose the template under “Submit sequencing reads using FAST5 file” option – fast5 here is just used as a blanket term to refer to all ONT upload.

Here’s what my filled out template for this particular data looks like – it’s just a TSV file.

FileType	OxfordNanopore_native	Read submission file type			
sample	study	instrument_model	library_name	library_source	library_selection	library_strategy	file_name	file_md5	design_description
SAMEA115768912	PRJEB76884	GridION	ONT_SQK-NBD114_a_sp_21022	GENOMIC	other	WGS	ont_r10_5kHz_dna_a_sp_21022_blow5.tar.gz	xxxxxxxxxxxxxxxxxxxx	SQK-NBD114 native barcoding kit with long fragment Buffer. 400Bps translocation speed

Once the sample is registered along with necessary metadata, it’s time for data submission. Methods for uploading the actual data for ENA roughly matches those of NCBI, so if you’re familiar with one there shouldn’t be any issues handling the other. For total beginners, I want to recommend avoiding the web/apsaras upload if possible, especially if you’re on a ‘nix system. FTP works, and it works well even for larger uploads.

ENA specific FTP upload method is described here: https://ena-docs.readthedocs.io/en/latest/submit/fileprep/upload.html#uploading-files-using-command-line-ftp-client

To summarize briefly – with ENA you can log in to a dedicated FTP upload space using your ENA login credentials. Using lftp from command line, it would look something like:

lftp -u yourENAusername ftp://webin.ebi.ac.uk

After which you’ll be prompted for your account user password. Once logged in you can simply upload your data right away – ENA already has the metadata TSV file at hand with the filename and MD5 checksum description.

mput ont_r10_5kHz_dna_a_sp_21022_blow5.tar.gz

Uploading from US East, I got about 20M/s, which isn’t too bad. Once the upload is finished, check “Run Files Report” under raw reads, which looks like:

It should show something similar to this:

And it’s done! We’ve successfully uploaded a third party format ONT long reads raw data to ENA for archiving and distribution. Much thanks again to Hasindu and Psy_Fer_ for telling me about this!


Now, you might be wondering what we could actually do with slow5/blow5 data outside of archiving with them and converting them to fast5/pod5.

I think the beauty of blow5 archive is how intuitive it is to access, understand and analyze them using conventional command line and other downstream tools. It’s TSV with header, so there’s not a whole lot of structure or schema to worry about if you’re already familiar with linux system and piping in general. Hasindu and the team has a wonderful page of bash one liners capable of carrying out some extensive analysis here: https://hasindu2008.github.io/slow5tools/oneliners.html

Here’s quick demonstration of what I was able to get going in about an hour or so of tinkering:

First, let’s note what each of the columns mean on a slow5 file (or blow5 accessed through slow5tools)

     1	        read_id
     2	read_group
     3	digitisation
     4	offset
     5	range
     6	sampling_rate
     7	        len_raw_signal
     8	raw_signal
     9	end_reason
    10	channel_number
    11	median_before
    12	read_number
    13	start_mux
    14	start_time

I’m going to take the length of the raw signals from column 7 and channel number from column 10 and make some quick plots. A bird’s eye view of length of individual raw signal events per channel might show us something interesting.

#let's isolate all signal length and channel events from the raw data file
slow5tools view all_b20_p.blow5 | grep -E -v '#|@' | cut -f 7,10 | sort -n -k1 > all_channel_length.tsv

And then let’s use a Julia environment for quick plotting fun.

#creating an isolated Julia environment - as is the norm these days, don't install packages into a base environment of a programming language
generate ont_aog

#aog will be doing the heavy lifting with CairoMakie backend
using AlgebraOfGraphics, CairoMakie, CSV, DataFramesMeta

#importing tsv
channel_data=CSV.read("all_channel_length.tsv", delim="\t", DataFrame)

#and going for a quick scatter plot
sca_plt = data(channel_data) * mapping(:Channel, :Length; color = :Channel) * visual(Scatter; markersize = 1)
with_theme(theme_black()) do
	save("channel_length_scatter_black.png", draw(sca_plt, scales(Color = (;colormap = :hawaii))); px_per_unit = 2)
	end

Running above on the slow5tools extracted raw signal data gives us:

Some impressive outlier signal read lengths here – 1713119 one must be the one on the top center-right portion. Unfortunately, it looks like we had some dead channels here – some are just dark from the get-go. Let’s split the 512 channels into 4 and see if there’s a noticeable signal length yield difference.

awk '$1 <=128' all_channel_length.tsv | datamash mean 2
35554.436621292

awk '$1 > 128 && $1 <= 256' all_channel_length.tsv | datamash mean 2
37821.437803739

awk '$1 > 256 && $1 <= 384' all_channel_length.tsv | datamash mean 2
38033.006384172

awk '$1 > 384 && $1 <= 512' all_channel_length.tsv | datamash mean 2
36449.37334595

The first and the last quadrant does not seem to fare very well. Let’s get a channel to read number plot as a comparison. That’s column 12 – which contains a number that counts upwards from zero every time a read is generated from the channel.

First graph using density model seems to show a smaller range of outperforming channels in read generation events – a linear block of them, in fact. The second plot shows certain channels simply doing better than the rest at generating read events (not all read events are actual reads, so we’re just looking at a data point here).

We can expand out the same method and plotting script to individual read files generated during the sequencing experiment. Here’s another sample of read event length against channel applied to just a single read file, not the whole dataset.

There’s certainly a broader range of activity up close – though the dead channels distribution here is more or less a mirror of the whole sequencing experiment plot we’ve seen earlier. Going through each of the read sets (including ‘fail’ category reads) with a fine toothed comb might net us some interesting observation.

With each of the signal-reads being huge matrices of numbers, there are endless variety of different tools and methods that could be used to visualize and experiment with them. And the slow5tools seem to be the fastest and most accessible way to crack open the raw signals for further research.

4th annual Halobacteria sequencing: genome extraction

It’s finally that time of the year – our annual Halobacteria mutant strain genome sequencing and assembly.

Quick recap; our lab has an acidophilic/tolerant (robust growth at pH 4.0~5.0) Halobacteria mutant strain from around 2019. I’ve been culturing and passaging them at fixed intervals since (based on their O.D rather than specific time, though it comes out to around once every two weeks), and have been sequencing the mutant strain once every year starting 2020. With falling price of Illumina sequencing, an annual 400Mbp short read coverage for our strain costs about $8 a month. Not necessarily cheap (for me) but certainly within the affordable range. I’m hoping a long-term continuation of this study will end up painting an interesting picture of how a microbe might (or might not!) gradually change within continuously passaging amateur biology research setting.

I decided to modify this year’s extraction protocol a bit. While extracting HMW DNA from Halobacteria is a relatively trivial process, our in-house protocol (available here: dx.doi.org/10.17504/protocols.io.ewov1qjx7gr2/v1) can sometimes have issues with purity due to occasional difficulty ensuring even mixing and suspension of the samples before and during RNase and proteinase K incubation step. A simple solution would be increasing initial suspension volume by adding 100ul dH2O to the 100ul Edward’s buffer, making the sample itself generally easier to mix and vortex properly without causing additional bubbles.

Based on the results, I think the new slightly modified protocol works far better than I expected.

>If using PCR machine for RNAse/proteinase incubation step, pre-heat the PCR machine. I normally use a program with 37C for 15 minutes, 55C for one hour, and 95C for 10 minutes. 

>Spin down 1ml of saturated Halobacteria culture for 1 minute at max speed (14k rcf on the lab microcentrifuge). Fully decant the supernatant.

>Resuspend the pellet with a mix of 1:1 Edward's buffer to dH2O (200ul total). The pellet is unlikely to resuspend just through pipetting - physical and vigorous mixture with the tip of the pipette is necessary. The cell mixture at this stage will look clumped with viscous texture.

>Carefully transfer 200ul of resuspended mixture to a PCR tube, ideally taking the sample from the bottom of the tube and avoiding any foams or bubbles. The sample will be viscous, so a slow and deliberate movement is necessary - wait for the sample to slowly move into the pipette tip, and then place the pipette tip at the very bottom and to the side of the PCR tube, touching one of the walls. Slowly and deliberately push out the mixture. Carefully moving up the pipette tip along with increasing level of the mixture in the tube can help with reducing bubble formation.

>Add 2ul of RNase A directly into the sample suspension - visually confirm all content was ejected from the pipette tip. A couple of pipetting in and out within the sample seems to help occasionally. Vortex in short bursts 10 times. If everything went well there will be zero to negligible amount of bubbles in the sample.

>Incubate at 37C for 15 minutes

>Add 2ul proteinase K directly into the sample suspension - visually confirm all content was ejected from the pipette tip. Pipette in and out a couple of times and vortex in short bursts 10 times.

>Incubate at 55C for one hour, and at 95C for 10 minutes to deactivate the proteinase
*If everything worked properly, the sample should look evenly mixed and completely clear, appearing transparent orange color for Halobacteria sample. If you get sedimentation or blobs of original sample, it might be worth restarting the process from beginning.

>Transfer the sample to 1.5ml tube, and estimate total volume of the sample.

>Add 1/10 volume 3M sodium acetate (NaOAc pH 5.4) - this is 1/10 volume of the eventual total mixture for this step. So you need to add -> initial sample volume from PCR tube (200ul) x 2 / 10 = 40ul
>Add 1:1 sample volume of 100% isopropanol (room temperature works better). If initial sample volume from PCR tube is 200ul, you'd add 200ul of 100% isopropanol
*This step can be confusing - normally it's easier to just add the 1:1 isopropanol, and then adding 1/10 3M sodium acetate. However, the very viscous nature of Halobacteria sample at this stage sometimes impedes proper mixing of the smaller volume NaOAc, the sample and 100% isopropanol. For long-read sequencing there's also a concern that too vigorous of a mixing process here can fragment the DNA.

>Invert the tubes to mix 20x

>Spin down at max speed for 5 minutes, decant fully using a pipettor
*If everything worked correctly up to this point, we get a nicely compacted, flat cell pellet at the bottom of the tube. If we have incomplete mixing/incubation we get a half-transparent ball of gel at the very bottom of the tube. If you have the latter, you could still proceed, but the output will be somewhat dirty and not suitable for pore-based sequencing processes.

>Resuspend with 1ml of 70% EtOH, and go through 3x washing steps (1x washing step: resuspend with EtOH, pipette up and down ~10x against the pellet, spin down at max for 5 minutes, decant fully)
*If everything worked, we're going to end up with a smaller, flat white pellet at the bottom of the tube with fuzzy periphery starting around 2nd wash. If your sample retains color, it usually means too much of the cellular contaminants made it over.

>Fully decant the sample using a pipettor. Dry the sample for about 5 minutes. Overdrying will plasticize the pellet - proper time frame is dependent on condition of the lab. Giving a quick smell test helps, we want to get rid of as much of the alcohol as possible.

>Resuspend in sterile dH2O, sometimes preheating the dH2O to 50~55C helps to dissolve the extracted pellet faster.

>Incubate the tubes with shaking at 37C for one hour. Placing an open plastic container (like pippette tip box lid) in the shaker and letting the sample tubes roll around on their own works.

>Incubate the sample at 4C overnight
*At this stage, a good extraction will look completely transparent. If you see coloration, do not use for nanopore based sequencer - you'll lose the flowcell.
The Halobacteria sample after 1 minute spin down. I keep on getting these multi-layered coloration.

Here’s how the Halobacteria pellet looks like after the first one minute spin down. Our culture in particular keeps on getting these multi-layered coloration, but the samples aren’t contaminated (16S screened). I wonder what they are.

Resuspended Halobacteria cultures prior to RNase/proteinase incubation.

This is how the sample looks immediately after resuspension in 200ul 1/2 Edwards buffer mix. Notice the clumping and the bubbles. Even these amounts of bubbles are drastically reduced from what we can get from a more concentrated 100ul suspension.

Samples after incubation with RNase and proteinase

Proper mixing and incubation of the samples should result in even, transparent mix like the ones in above picture. If you notice clumping at this stage, it’s well worth it to simply start the process again.

Spin down pellet after the incubation step. Bottom tube is the properly mixed sample. Top tube is indicative of incomplete incubation

Above picture shows the major difference between a properly mixed and incubated sample versus an incomplete one. The bottom tube’s pellet is flat with fuzzy ends – this is how it’s supposed to look like. The top tube’s pellet is round, retaining 3D structure. It’s gelatinous and sticks to pipette tips. Too much of cellular detritus made it to the top sample.

Two pictures of how each of the above samples look after EtOH washing step (3x). The picture on the left is the final washed pellet from the contaminated sample with round, gel like pellet. The picture on the right is the final washed pellet from the properly incubated sample with flat, fuzzy pellet. Notice how the properly incubated sample lost its color.

Gel image of the extracted samples. NEB 1kb extend ladders on the sides. From left to right – 1A 2A samples using previous extraction protocol. 1B 2B samples using new extraction protocol

As usual, here’s a quick gel electrophoresis picture to make sure we’re getting good HMW extraction from the process. The ladders on the sides are NEB 1kb extend, with top band weighing in at 48.5kb. The two lanes on the left are control samples with 100ul resuspension. The two lanes on the right are test samples with the new protocol with 200ul resuspension. They essentially look similar – this step is only useful to let us know that we have HMW DNA in the final extraction.

A proper nanodrop result give us a more comprehensive look at the differences. The ‘a1’ sample on the left is from the older style genome extraction protocol. The ‘b1’ sample on the right is from the new 200ul suspension genome extraction protocol. ‘a1’ actually is on the somewhat cleaner side of some of these older protocol genome extracts – but ‘b1’ is simply a far better genome extract here.

All in all, the results look nice – I’m strongly considering spending some extra money and sending this sample out for a full short & long read sequencing for the fourth year run. It should be enough to let me assemble a fully closed genome of the particular strain.

Before I go – here’s a quick snap of a visitor that came around here somewhat recently.

Wrapping up 2023 – review as an amateur biologist

It’s six in the evening where I am – sitting aside a window looking out into a field of almost complete darkness. Winter evenings set in faster and darker than ever, and I’m constantly reminded of a small part of my brain still stuck in late summer and its long, drawn out fading of the day into the night. I’m pretty sure that part of my brain still thinks it’s 2021 too.

These blank, ponderous moments in time always had a mysterious power over me, turning me inward for introspection. This is the last evening of the year 2023 – everything that has been done is done and will never be revisited or revised. How was 2023 for a loading dock worker who fancies himself an amateur scientist-in-making?

Life can be tough, and I’ve certainly had my share of hurdles – however, looking back I think my 2023 was defined by a great sense of gratitude and camaraderie from my fellow human beings.

I joined and quit a biotech company job within the span of a couple of months, followed immediately by a part-time research gig with a professor from Canada with perhaps one of the most interesting research scope I’ve ever had the luck to encounter in person… This job provided with me with a great, rare opportunity to definitively confirm to myself that I am not ready for such a position.

I should be sharper and more focused in my questions, more diligent in my observations, more stringent in my reviews and far, far more tenacious in pursuit of truth with the least amount of compromise.

In the end I couldn’t bring myself to charge him for subpar output, and we parted ways as such. Hopefully there will be an opportunity in the future for me to address the sense of debt and gratitude I still hold for being offered the opportunity in the first place.

Life playing bit of a merry-go-round, a couple of months later I was hired by another professor I’ve collaborated with before, as a genuine research scientist, albeit part-time. Luckily this time the research scope fell within more familiar realms of microbiology and sequencing, where the brunt of study still fell within the ‘librarian of genomes’ role I felt comfortable with. Still – I’m not so naive to assume a position like this was offered to a loading dock worker without any internal convincing and risk taking.

The research scientist role took the lion’s share of my time and energy through most of the year, with me trying to reflect some of the lessons from the earlier attempt at academic research. In practice this meant more stringent literature reviews, better planning of deliverables in each stage of the broader research goal, and planning out my schedule so that I can attempt to learn new skills and tools, rather than only reusing what I knew how to do at that moment in time.

The research, focusing primarily on phage genomics and some phylogeny, is still ongoing – we’ll likely only see real sharable results going into Spring of 2024. Hopefully we’ll be able to make something of this, though the primary thesis we started out with looks a little tenuous based on new data.

As an amateur biologist, running thread through major events of the year had been pursuit of depth. It’s Leeuwenhoek all over again – I think I can grind my own lens and see the animalcules, the kind no one’s seen before. I might even be able to make out certain patterns of behavior, hinting at some sort of truth in nature. But the final push – something that crosses the threshold to pull a well-founded suspicion to a certainty (within means) is still lacking.

Finding my own, amateur biologist way to somehow cross this threshold will define the coming new year, 2024.

And with that, Happy New Year to everyone – I hope everyone finds peace and vigor enough to pursue their dreams in the new year. And I hope we can all come together to endeavor to create beauty in this world, not merely try to possess what’s already there.

Treebuilding common E. coli strains

As people say, time flies. Unbearable heat of the summer at the warehouse lab, compounded by emanations from overheating computers are still vivid in memory as if it was yesterday. Now, I’m looking at the last of autumn leaves as the flora sheds their colors and turn gaunt one by one. We’ll be welcoming winter in the next few weeks at most.

I recently took a small sojourn to a friend’s place, hoping to recover from a recent health issue. Over one of our usual conversations on labs and amateur biology he requested an interesting review – we’re well aware of standard E. coli strains available to many amateurs for workhorse cloning and protein expression, but how distant are they from one another? Are they all ‘just E. coli’ as some observers quipped in the past, readily swappable from one strain to another as need arises?

This was a good question, a perennial encore in casual lab chatters that everyone seems just a little too busy to look up at the time. A perfect fit for a long evening’s search, gauging a broader conceptual distance and likeness between some of the more common workhorse E. coli strains.

I won’t be addressing any specific technical differences between the strains here. If you’re looking for specific genotype differences between well known E. coli strains, https://openwetware.org/wiki/E._coli_genotypes is a good starting point. For now, this is a quick note on broader genome to genome comparison between a small number of commonly available E. coli strains, to act as a conceptual tool to help me get a perspective.

Associating the E. coli strain names to available assemblies presented some initial difficulties. For future reference, looking up common names in https://bacdive.dsmz.de/ usually led to proper ATCC/DSM/etc numbers or strain types registered on Genbank. I recorded NCBI taxon IDs for each of the assemblies when available, though there were some curious cases here and there. Like how Takara’s BL21 strain is registered under general E. coli taxon ID while NEB’s strains are registered with BL21 specific ID.

W3104 assembly recorded on the table below is currently suppressed for contamination issues, and was left out of all subsequent analysis. Hopefully it’ll be corrected soon. For whatever it’s worth, the particular strain is supposed to be derived from W3100, which is DSM 30083/ATCC 11775, also included in the below table as strain MM294. I’d assume they’re phylogenetically closely related.

NEB Turbotaxid:562https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_013167055.1/
DH5alphataxid:562https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_022221385.1/
NEB 10betataxid:562https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_013167035.1/
K-12taxid:511145https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_000005845.2/
Btaxid:37762https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_001559635.1/
Btaxid:37762https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_001559615.2/
BL21 (NEB)taxid:511693https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_013166975.1/
BL21 (NEB DE3)taxid:469008https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_013167015.1/
BL21(TaKaRa)taxid:562https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_000833145.1/
Ctaxid:498388https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_002079225.1/
C600taxid:866789https://genomes.atcc.org/genomes/377ef90c440d444d
MM294taxid:866789https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_003697165.2/
OP50taxid:637912https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_004355015.1/
W3104taxid:562https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_018430865.1/
JM109taxid:562https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_028743375.1/
HB101taxid:634468https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_026228125.1/
E. fergusonii*taxid:564https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_020097475.1/
NCTC86EC**taxid:562https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_900092615.1/
List of common lab E. coli strains. * is included as an outgroup ** not a common lab strain – but first E. coli described by Theodor Escherich, namesake of Escherichia

Search for E. coli genomes associated with the common strain names led me to two fascinating papers. First, a sequencing paper on E. coli strain NCTC86EC, likely the very first E. coli strain described by Theodor Escherich in 1885, and a follow up paper elucidating evolutionary journey of the particular strain as it was archived and passaged from lab to lab over a period of 133 years. I decided to include the genome of the strain in the analysis, along with E. fergusonii as root.

Dunne KA, Chaudhuri RR, Rossiter AE, Beriotto I, Browning DF, Squire D, Cunningham AF, Cole JA, Loman N, Henderson IR. Sequencing a piece of history: complete genome sequence of the original Escherichia coli strain. Microbial genomics. 2017 Mar;3(3).

 doi: 10.1099/mgen.0.000106
Desroches M, Royer G, Roche D, Mercier-Darty M, Vallenet D, Médigue C, Bastard K, Rodriguez C, Clermont O, Denamur E, Decousser JW. The Odyssey of the ancestral Escherich strain through culture collections: an example of allopatric diversification. Msphere. 2018 Feb 28;3(1):10-128.

doi: 10.1128/mSphere.00553-17

Initial phylogenetic tree was built with my favorite tree building suite, GToTree (https://github.com/AstrobioMike/GToTree) using its built-in Gammaproteobacteria SCG HMM set of 172 genes. One caveat to note – this version was built using default FastTree program, which might not be an optimal idea for comparing lots of long similar sequences containing inversions and shuffling. I’ll use the generated data to build a proper maximum likelihood tree later.

Rooted phylogenetic tree of E. coli strains from above table

Clustering of E. coli genomes here is pretty self-evident. MM294 remains closest to the outgroup E. fergusonii. E. coli B strains cluster with OP50 and BL21s. E. coli HB101/K-12/JM109/DH5a cluster mostly together, along with other specialty strains from NEB. The 133 year old E. coli NCTC86EC is on its own branch between the two larger clusters, along with E. coli C.

ANI cluster map of the same genomes

ANI cluster map created using FastANI (https://github.com/ParBLiSS/FastANI) & UPGMA clustering above follows the pattern we’ve seen on the phylogenetic tree. All genomes are, of course, extremely similar, though I’m sure we can coax more differences by playing around with kmer settings.

FastANI has a built in visualization option that produces 12 column file containing coordinate to coordinate ANI percentage matches. The text file can be run through various visualization tools (in fact, I’m working on one myself). In this case I used a built-in script for PyGenomeViz package (https://github.com/moshi4/pyGenomeViz) to create genome to genome ANI graphs with the original NCTC86EC strain serving as reference.

These are comparisons of the genomes of same species, yet I see more differences among them than expected. Inversions and shufflings within roughly recurring regions of the genomes is also clear to see, though OP50 genome contigs probably should have been scaffolded/moved prior to ANI comparison and visualization. I’m working on something that adds a gbk tick to above graphs to more comprehensively show which regions of the genomes are inverted/linked against another – mostly working within phage context, but that’s a story for another time.

All in all, I think this presents an interesting overview – and a small food for thought for our undegraduate and amateur biology friends when thinking about genomics among strains of the same species. There are certainly noticeable differences even within our human-made taxonomic distinctions!

Rough translation – the history of academic poster

I had an idea to look into the history of academic posters while working on the Nanopore Community Meeting 2022 material – almost a whole year ago now.

It turns out the academic poster is a very recent concept (relatively speaking), with original introduction traced back to 1967, and the type of large, colorful printouts we’re all familiar with becoming the norm only after the 70’s (mostly on account of expenses – physical printing technology did advance significantly between then and now).

A commonly referenced paper addressing the origin of the academic poster (among other things, like references to best practices and future of poster in a more digital age) is “Lire debout. Le poster comme pratique de lecture dans le monde scientifique” written by Françoise Waquet, available here: https://teca.unibo.it/article/view/13137

Waquet F. Lire debout. Le poster comme pratique de lecture dans le monde scientifique. TECA. 2012 Mar 1;2(1):9-22. DOI: https://doi.org/10.6092/issn.2240-3604/13137 

And the AIP history newsletter from 2008 that led me to the author’s paper is located here: https://repository.aip.org/islandora/object/nbla%3A275060#page/5/mode/1up

Here’s a (very badly) translated excerpt on the origin of the academic poster by yours truly:

The history of the poster is one of widespread and rapid success in scientific communication. The "first" posters or very similar products were presented as early as 1967 at the Carshalton Medical Research Center, a major British research institute. These were “prepared cards” on which researchers compiled results, mostly graphs and tables – the material that was usually shown through slides; the sheets were placed on panels and the authors stood beside them for a discussion. This type of presentation, born in the field of biochemistry, won through the whole discipline, then in the years 1975 and following, was introduced in the international and national conferences of all life sciences.(#9) Once adopted, utility of the posters were almost never questioned; they largely replaced oral communications.

The history of the poster is one of widespread and rapid success in scientific communication. The "first" posters or very similar products were presented as early as 1967 at the Carshalton Medical Research Center, a major British research institute. These were “prepared cards” on which researchers compiled results, mostly graphs and tables – the material that was usually shown through slides; the sheets were placed on panels and the authors stood beside them for a discussion. This type of presentation, born in the field of biochemistry, won through the whole discipline, then in the years 1975 and following, was introduced in the international and national conferences of all life sciences.(#9) Once adopted, the posters were almost never questioned; they largely replaced oral communications.

The author suggests the practice of ‘prepared cards’ for rapid (and perhaps less formal) presentation of ideas is something that happened at the right place at the right time – that late 60’s and onward saw a huge explosion of academic research population and programs, and quickly necessitated modes of presentation that could save on time and space.

It’s easy to consider academic research as separate from the ebb and flow of common life, but it looks like at least the poster tradition is a product of post-war boom generation and increasingly active and crowded venues for exchange of ideas.

If the common threads of academic medium and its traditions adapt to changes of society and culture, what other new developments and transformations await us now? I think the 20 year long rumbling we’ve been observing in the field of academic publication is a ripe candidate for a new, practical experiments – with the caveat that contemporary scientists must now deal with billion dollar companies actively promoting status quo. Perhaps this too is part of the shift in society and culture exerting its influence on how science is developed and perceived.

If you want to check out the rest of the paper, I have a terribly translated (but readable) version in English at https://github.com/naturepoker/lire_debout-2012-translation

Halobacteria mutant sequencing; 3 years in

It’s a special week over at Binomica Labs.

As mentioned in some of my previous notes here, I’ve been maintaining a continuous culture of acid-tolerant Halobacteria NRC-1 line for the past couple of years, with short read sequencing every anniversary of initial inoculation date – Feb 14th 2020 (and yes, the pandemic was definitely a factor in starting this study).

Halobacteria NRC-1 has one of the better reference genomes, remarkably uploaded in 2001 before much of modern NGS tools became the norm (the reference genome is available here, GCF_000006805.1). The goal of this small ongoing observation is to continuously grow our Halobacteria culture long-term and record any shifts and variants against the reference, using laboratory resources that might be representative of more involved amateur biologists.

It’s important to note that any observation made in this fashion would be a happenstance snapshot of a very stochastic system. Whatever the change we might observe in a study like this would be representative of only this particular instance of culture and sequencing run (if even that, as this very fascinating study suggests: https://doi.org/10.1038/nature24287). So I would be hesitant to call this a research on evolutionary biology, or even mutation accumulation – at current scale this would merely be series of interesting observations that could one day result in a proper research. A microbiological equivalent of bird-watching, I think.

Despite my preference for long-read sequencing in situ (such as with ONT), this study relies on third party short read sequencing. At this time it’s difficult to compete against pricing from third party sequencers with economies of scale on their side, though hopefully new flongle studies planned for later half of this year will change things. I’m expecting about 650mb data throughput at the cost of $110 USD – not cheap, but certainly not insurmountable if taking place only once a year.

Based on published Halobacteria doubling time (1.86~/h) and OD measurement in our lab, current sequencing run is for a 44595th generation of Halobacteria NRC-1 (HaloM_11) in continuous acidic (pH 4<) media culture. Translated to human generations (with 25 years per generation) this is similar to almost a million (1,114,883 years) years of human propagation – though this is more of a literary interpretation of the two organisms’ life time.

Standard Binomica Labs whole genome extraction method was used, as described below:

>Spin down 1 ml of sample for 1 minute at max speed and decant
>Resuspend vigorously with 100ul Edward's buffer
>Transfer carefully to PCR tube - mixture will be viscous
>Add 2ul of RNase A and mix vigorously, vortex for 10 seconds
>Incubate at 37C for 15 minutes
>Add 2ul of Proteinase K and mix vigorously, vortex for 10 seconds
>Incubate at 55C for 1 hour, and deactivate via incubation at 95C for 10 minutes
>The mixture should be clear and evenly mixed at this point
>Transfer to 1.5ml eppendorf tube
>Add 10% 3M (pH 5.4) sodium acetate, and 1:1 volume of 100% isopropyl alcohol 
>Invert tube 10 times - precipitates should begin to form
>Spin down at max speed for 5 minutes (if you can observe semi-clear, orange mixture at this point, RNase and Proteinase mixing and incubation did not work properly)
>Decant carefully
>Add 1ml of 70% EtOH and resuspend the pellet
>Spin down at max speed for 5 minutes
>Repeat 70% EtOH resuspension and washing step at least 2 more times
>Decant completely and dry the pellet for 5 minutes - do not let the pellet overdry
>Resuspend in storage buffer of choice or dH2O to wanted volume (I do dH2O up to 200ul)
>Incubate in a 37C shaker for 1 hour - the extract is likely to be not dissolved fully
>Incubate in 4C overnight - depending on yield, the extract will be extremely viscous
Halo mutant whole genome extraction on a gel, 1 and 2 from same culture. Ladders on the sides are NEB 1kb extend, with top ladder size of 48.5kb. I ended up sending out sample #2

Gel electrophoresis image of the genome extracts is attached above. The last two lanes are Vibrio genome extracts for a friend’s project. Sample #2 was sent out for sequencing – with results expected around mid-March.

Mapping the reads against 2021 sequencing reads and the reference genome should provide for an interesting observation. I’m hoping to document the observation and methods on our upcoming amateur biology zine – Willowlands.

Edit Jun 12th 2023: protocol described here is also available on protocols.io with citable DOI:

DOI: dx.doi.org/10.17504/protocols.io.ewov1qjx7gr2/v1

As a bonus, here’s a look at our extremely reliable tube tumbler used for long-term cultures in the lab. It’s been running nonstop for close to 4 years now:

MERRYCHRISTMASANDHAPPYNEWYEAR

Alphabet representation for protein amino acid residues can fun to play around with – let’s see what a season’s greeting looks like with a quick blast search:

>greetings
MERRYCHRISTMASANDHAPPYNEWYEAR

Dumping above into NCBI blast shows:

Interesting! I’m surprised we got any hits at all – let’s look at the alignment quality:

Ah – this is no good. Essentially we have a short fragment crossover resembling below at best:

HEPPYEEWFE
HEPPYPEWYE
HAQTYNEWYEA

Welp, it goes to show, simply getting a list of alignments off a database sometimes doesn’t mean anything.

Hope everyone’s having a wonderful holiday break and wishing you a HAQTYNEWYEA!