Samuel Galvão Elias
Curriculum Vitae
0000-0001-9138-8845
LepistaBioinformatics
sgelias
lepista.com.br / lepista.io
Summaries
In Microbiology
With approximately twelve years of experience in Microbiology and Mycology, I have developed expertise across two major areas: Health Sciences (where my career began) and Life Sciences (the area of greatest depth and current focus). My trajectory in Mycology spans eight years of continuous research conducted in collaboration with leading Brazilian scientists, including Aristóteles Góes-Neto [1], Elisandro Ricardo Drechsler-Santos [2], João Paulo Machado de Araújo [3], José Carmine Dianese [4], Maria Alice Neves [5] and Robert W. Barreto [6] — all mycologists of national and international distinction. Throughout this period, I also established research partnerships with international scientists, among them Charles W. Barnes [7], Mary Catherine Aime [8], Harry C. Evans [9] and Dominik Begerow [10], all of whom I regard as direct and indirect mentors in the field. My primary areas of scientific interest encompass the genetics, taxonomy, ecology and evolution of microorganisms, with emphasis on wild fungi and bacteria lacking commercial relevance, as well as pathogens of agricultural and environmental concern.
Complementing this background, I hold in-depth knowledge of biostatistical and bioinformatic analytical techniques (see section Bioinformatics), with a particular focus on genomic and metagenomic data analysis.
Driven by scientific inquiry and the elucidation of natural phenomena, a comprehensive account of my scientific contributions is available in the Publications section.
And in Bioinformatics
A self-taught bioinformatician with over ten years of experience in biological data analysis through biostatistics and bioinformatics methodologies. In recent years, professional focus has shifted towards the design and development of high-performance and high-availability bioinformatics tools tailored for modern, highly connected web environments.
Currently, core activities encompass the modeling, architecture, development and coordination of multidisciplinary technology teams dedicated to building high-performance biological data processing environments. Areas of technical expertise include solutions for parallel processing — a fundamental requirement for bioinformatics analyses at industrial scale — concurrent processing for distributed computing environments, and the automation of large-scale pipelines.
Additional contributions include the development of web-native and public domain tools (see section Public Tools), among them Mycelium API Gateway and Mycelium WebApp, a pair of solutions focused on distributed access control for complex API environments; Blutils CLI and Blutils UI, a pair of tools dedicated to biological sequence analysis; and Classeq, a biological sequence classifier for post-phylogenetic placement. All aforementioned tools are made freely available as open-source projects, accessible and auditable through the official GitHub of the Lepista Bioinformatics organisation.
Links:
[1] A. Góes-Neto; [2] E. R. Deschsler-Santos; [3] J. P. M. Araújo; [4] J. C. Dianese; [5] M. A. Neves; [6] R. W. Barreto; [7] C. W. Barnes; [8] M. C. Aime; [9] H. C. Evans; [10] D. Begerow.
Public Domain Software
Over the past three years, independent work has been dedicated to the development of bioinformatics and web-native tools released as free and open-source solutions. The overarching objective is to deliver meaningful contributions to the scientific community and to society at large, through tools of practical relevance to both specialist and non-specialist users. The principal tools developed include:
Zombie Crab Project
An open-source platform that gives every user their own real, isolated AI agent behind a single authenticated entry point, exposed through an OpenAI-compatible HTTP API. It was built because self-hosted assistants such as PicoClaw follow a “one agent, one owner” model: when several people share the same process, a single prompt injection or leaky tool is enough for one user to read another’s conversations, files and secrets. The stack answers this with three independent layers — an edge (the Mycelium gateway), which authenticates the caller and injects a verified, unforgeable identity; an orchestrator (crab-shell-proxy), which gives each (agent, user) pair its own non-root container and volume; and a swappable agent harness — so that isolation is enforced by the kernel rather than by application code. Agents can scale to zero when idle or run continuously, and the chat client adds workspace memory, a knowledge graph, scheduled tasks, files and secrets management; the whole stack is deployed with Docker Compose and licensed under MIT or Apache-2.0. It is the choice for any organisation that wants to offer AI agents to many people at once without letting one user’s agent ever reach another user’s data.
crab-shell-proxy
The Go orchestrator at the centre of the Zombie Crab stack. For every request, it reads the target agent and the caller’s account from the identity injected by the gateway — always the stable account id, never the e-mail — ensures that this member’s own container is running, starting it on demand and stopping it when idle, and relays the conversation to it. It translates OpenAI-style HTTP into each harness’s native protocol and answers explicitly, naming the harness, whenever a capability is not available, instead of failing silently. It also hosts the model registry, the knowledge-graph memory (served to agents over MCP with scoped tokens), scheduled tasks and projects, and it talks to Docker through its own minimal client, keeping secrets out of the images. You would use it whenever each user must get a dedicated, disposable agent environment on demand, managed by a single, auditable control plane.
crab-exoskeleton-webapp
The chat web application of the Zombie Crab stack, through which members talk to their agents and operators govern the fleet. Built with Next.js 15 as a backend-for-frontend, it keeps every credential on the server: the browser holds only an HTTP-only session cookie, with no token, account id or upstream address, and each request flows through the gateway and the orchestrator before reaching the agent. Sign-in is passwordless, through a magic link with a six-digit code; conversations are indexed in PostgreSQL; and the configuration is read at request time, so a single image serves every deployment. It is backed by a suite of more than 1,700 automated tests. You would use it to give end users a familiar chat experience on top of isolated agents without ever exposing tokens or infrastructure details to the browser.
crab-ganglion-harness
The agent runtime written for the Zombie Crab project and now its default harness. Running inside each member’s container, it holds the conversation, calls the language model, executes tools and records the transcript, exposing a native HTTP+SSE interface through which narration and reasoning are streamed live. It was created when the limits of the original assistant began to cost more than they saved, and it is designed for safety and simplicity: a single static Go binary with no third-party dependencies, whose shell tool runs inside a Linux Landlock sandbox that confines each turn to its own workspace and refuses to start if that protection is unavailable. It ships with web search and fetch, image generation, sub-agents, history search and MCP tools, caps the number of tool iterations per turn, and follows a hexagonal architecture enforced by tests. You would use it when you need an agent runtime that is small enough to audit and whose sandbox is guaranteed by the kernel itself.
harness-sphere
A single-binary OpenTelemetry watcher, written in Rust, that turns the state of the host, the gateway, the orchestrator, the web application and every agent container into standard metrics. It was built because the stack previously emitted no telemetry at all, making it impossible to tell whether an agent slowed down because of the agent or because of the machine underneath it. It models six explicit layers, collects data per member, and is designed never to take itself down: each collector runs in isolation, failures are contained and retried with back-off, and it deliberately never touches the Docker socket, token costs or conversation content. You would use it to observe a multi-tenant agent platform in production with Grafana or any OpenTelemetry backend, without compromising the privacy of users’ conversations.
crab-mangrove-network
An experimental federated memory network that lets agents share knowledge over the ActivityPub vocabulary, each agent being an identified bot actor owned by its human and governed by Mycelium roles. It addresses a limitation of private memory: two people working on the same problem build two disjoint knowledge graphs and rediscover the same facts twice. Authority flows in one direction only — nothing can be shared beyond the sharer’s reach, agents cannot broadcast to groups, and every share waits for the recipient’s human to admit it — while memory is kept as an append-only log of ed25519-signed activities, so that one author can never overwrite another’s. Written using only the Go standard library, it is entirely optional, and the stack runs unchanged without it. You would use it when teams want their agents to build on each other’s findings while every human keeps control over what reaches their own agent.
Mycelium API Gateway
An open and free API gateway, written in Rust, designed for modern, multi-tenant and API-oriented environments. It centralises authentication, identity normalisation, routing and policy enforcement: coarse, role-based checks are made at the gateway, while a verified profile injected into each request — treated as an active capability object rather than a mere identity payload — allows downstream services to take fine-grained, contextual decisions. It supports tenants, account types and security groups, sign-in through magic links, OAuth2 providers and Telegram, a JSON-RPC administration interface, webhooks, envelope encryption and key rotation, and it can expose downstream APIs as tools for AI agents through the Model Context Protocol (MCP). It runs with PostgreSQL, Redis and Vault or as a standalone binary with no external dependencies, and holds the OpenSSF Best Practices badge. You would use it to protect a set of APIs shared by several organisations behind a single entry point, keeping the security rules in one auditable place instead of scattering them across every service.
Mycelium WebApp
The official web interface of the Mycelium API Gateway, designed to allow rapid implementation of access policies to the gateway’s endpoints without the need for additional implementation. Built with React 19, Vite and TypeScript, it offers login through any OAuth 2.0 provider, a dashboard, and the management of tenants, accounts and sharing, roles, fine-grained connection strings and webhooks, with a mobile-friendly and multilingual layout. You would use it to administer a Mycelium deployment visually, delegating day-to-day access management to operators who do not need to handle the gateway’s configuration directly.
Blutils CLI
A high-performance command-line wrapper for NCBI BLASTn, written in Rust and distributed on crates.io, that improves on BLAST’s native parallelism and adds an exclusive consensus algorithm for taxonomic identification. When a query sequence matches several reference organisms with similar scores, Blutils resolves these multiple identities into a single consensus, using presets tailored to fungi and eukaryotes (ITS cut-offs), to bacteria (16S rRNA) or custom thresholds, and two strategies — cautious, which keeps the shortest reliable taxonomic path, and relaxed, which keeps the longest. It builds its reference database from the NCBI taxonomy and a BLAST database, and writes results as JSON, JSONL or YAML, ready for downstream pipelines. You would use it to classify large sets of amplicon or metabarcoding sequences quickly and with reproducible, explainable taxonomic assignments.
Blutils UI
A visual explorer for the consensus results produced by Blutils CLI, running entirely in the browser and published on GitHub Pages. The user simply loads the results file and explores it in three complementary views — a filterable table, a tree grouped by taxonomic rank and a network grouped by query sequence — together with a composition chart that explains why each consensus identity was selected. Because all processing happens on the client, the data never leaves the user’s computer. You would use it to inspect and communicate taxonomic results without installing anything and without uploading sensitive data to a server.
Classeq
An alignment-free phylogenetic placer of biological sequences, written in Rust, that positions new DNA sequences on a predefined reference phylogeny based on their k-mer composition. It is the improved implementation of the classifier developed during my doctoral research, and it avoids the cost of multiple sequence alignment while keeping the results anchored in an explicit evolutionary tree, supplied as a rooted Newick file with its reference sequences. It is available as a command-line tool and as an API server for distributed placement, and it emits OpenTelemetry logs and traces; the project is still under active development. You would use it to classify sequences quickly against a curated phylogeny, obtaining placements that can be interpreted in evolutionary terms rather than as simple similarity hits.
Private Software Registrations
Current professional focus is concentrated on the development of bioinformatics tools for the Agrobiotechnology sector, encompassing solutions for genomic and metagenomic data analysis as well as biological sequence processing. Parallel efforts have been directed towards the design and delivery of web-native, high-availability interfaces supporting software solutions for Brazilian agribusiness.
The registrations listed below, filed with the Brazilian Patent and Trademark Office (INPI) between 2022 and 2026, span metagenomic and genomic processing pipelines through to high-availability web applications for agribusiness.
2025
RobsonsAutomationHub
- Kind (Language): Automation (Python)
- Application: BR 51 2026 000320 2
AgroPortal
- Kind (Language): WEB (Typescript/HTML/CSS)
- Application: BR 51 2026 000317 2
SupernovaFungiDeep
- Kind (Language): Pipeline (Python)
- Application: BR 51 2026 000316 4
SupernovaBacteriaShallow
- Kind (Language): Pipeline (Python)
- Application: BR 51 2026 000315 6
SupernovaBacteriaDeep
- Kind (Language): Pipeline (Python)
- Application: BR 51 2026 000314 8
SilentBio
- Kind (Language): Pipeline (Python)
- Application: BR 51 2026 000312 1
ReportBuilder
- Kind (Language): Pipeline (Python)
- Application: BR 51 2026 000311 3
NaturalStackAPI
- Kind (Language): API (Python)
- Application: BR 51 2026 000310 5
EVAAuth
- Kind (Language): API (Python)
- Application: BR 51 2026 000309 1
CustomersAPI
- Kind (Language): API (Python)
- Application: BR 51 2026 000308 3
AgroPortalCustomers
- Kind (Language): WEB (Typescript/HTML/CSS)
- Application: BR 51 2026 000153 6
AgrobiotaSDK
- Kind (Language): LIBRARY (Python)
- Application: BR 51 2026 000152 8
2023
QuorumSensing
- Kind (Language): API/ETL (Python)
- Application: BR 51 2023 003328 6
PathFinder
- Kind (Language): Pipeline (Python)
- Application: BR 51 2023 003327 8
MetaDaVinci
- Kind (Language): Pipeline (Python)
- Application: BR 51 2023 003324 3
BioTax
- Kind (Language): API (Rust)
- Application: BR 51 2023 003321 9
BioReferee
- Kind (Language): API (Python)
- Application: BR 51 2023 003311 1
BioArchival
- Kind (Language): API (Rust)
- Application: BR 51 2023 003309 0
AgroReporterUI
- Kind (Language): WEB (Typescript/HTML/CSS)
- Application: BR 51 2023 003305 7
Agroportal
- Kind (Language): WEB (Typescript/HTML/CSS)
- Application: BR 51 2023 003302 2
AgroIndexAPI
- Kind (Language): API (Python)
- Application: BR 51 2023 003300 6
AgroBase-Rust
- Kind (Language): LIBRARY (Rust)
- Application: BR 51 2023 003288 3
AgroBase-Python
- Kind (Language): LIBRARY (Python)
- Application: BR 51 2023 003286 7
2022
QuorumSensing
- Kind (Language): API/ETL (Python)
- Application: BR 51 2022 003473 5
AgroBase-Python (update)
- Kind (Language): API (Python)
- Application: BR 51 2022 003472 7
AgroBase-Python
- Kind (Language): API (Python)
- Application: BR 51 2022 003471 9
Titration
PhD in Microbiology
The doctoral thesis, entitled “Fungi in Cerrado plants: an integrative approach”, resulted in the development of two bioinformatics tools: “GeneConnector” and “Classeq”. The former was designed to consolidate genetic information by leveraging metadata intrinsic to GenBank records. The latter, “Classeq”, is a biological sequence classifier that employs pre-built phylogenies to position DNA sequences with high performance, reduced computational cost and low error rate, through an API-accessible environment designed for machine integration.
Master in Fungi, Algae and Plant Biology
The master’s dissertation comprised a biogeographic study of the fungus Phellinotus piptadeniae, a species widely distributed across the neotropical region in association with legume hosts. Environmental modelling tools were applied, drawing on the fungus–host relationship, to infer the probable geographic distribution of the species. The study further identified a putative species complex within Phellinotus, characterised by a diffuse evolutionary history shaped by climatic fluctuations and ecological interactions.
Bachelor in Biological Sciences
Undergraduate training provided a solid foundation in pathogen–host interaction biology, with research conducted using bacteria and non-human animal models in the context of neurodegenerative diseases, including meningitis. This period also established core competencies in biostatistics and n-dimensional modelling, alongside the opportunity to collaborate with neuroscientists Tatiane Barichello and João Quevedo, both of international standing in their respective fields.
Publications in Journals
The scientific trajectory presented herein is defined by two major professional phases: an initial period dedicated to Health Sciences research, followed by a transition and subsequent consolidation in Life Sciences and Bioinformatics. The contributions listed in the sections below are organised by field of knowledge and presented in reverse chronological order.
Life Sciences and Bioinformatics
The following entries comprise scientific contributions within the fields of Life Sciences and Bioinformatics, covering publications from 2018 to 2024.
GeneConnector: Unlocking the full potential of Genbank metadata
Neotropical Studies on Hymenochaetaceae: Unveiling the Diversity and Endemicity of Phellinotus
Reinstatement and phylogenetic allocation of the palm rust genus Cerradoa in the Pucciniaceae, and establishment of Pseudocerradoa, gen. nov
Phytophthora theobromicola sp. nov.: A New Species Causing Black Pod Disease on Cacao in Brazil
Studies on the biogeography of Phellinotus piptadeniae (Hymenochaetales, Basidiomycota): Expanding the knowledge on its distribution and clarifying hosts relationships
Moniliophthora perniciosa, the mushroom causing witches’s broom disease of cacao: Insights into its taxonomy, ecology and host range in Brazil
A new section, Lactifluus section Neotropicus (Russulaceae), and two new Lactifluus species from the Atlantic Forest, Brazil
Phaeochorellaceae, Diaporthales: a new fungal family and a re-appraisal of Phaeochorella species.
Phylogenetic Relationships of Phaeochorella Parinarii and Recognition of a New Family, Phaeochorellaceae (Diaporthales)
Taxonomy, phylogeny, and divergence time estimation for Apiosphaeria guaranitica, a neotropical parasite on bignoniaceous hosts.
Crossopsorella, a new tropical genus of rust fungi
Health Sciences
The following entries comprise scientific contributions within the field of Health Sciences, covering publications from 2013 to 2014.