Run the CodeLaser engine yourself.
Some teams would rather drive the engine than hand the work over. We license it, install it on your side, and stay close while your team gets to work. It is the same engine we use, running the same plans.
What the engine is
It reads your project like a compiler, then keeps going.
An AST gives you the shape of the text. The engine resolves every type and tracks every reference the way a compiler front end does, then carries on past where a compiler stops: what can stand in for what, what is reachable from where, how data moves between fields and methods, and what else has to change when you change one thing.
The code model
Types and their hierarchies, members, every reference between them, reachability, data flow, mutability. Built once from your sources and the classpath, then kept current as the code changes.
The queries
Read primitives over that model: types, members, usages, call chains, package cycles, coupling, the standard metrics. Answers come back as data for the next step.
The operations
Move, split, extract an interface, retype a declaration, remove dead code and its cascade. One operation makes every edit it implies, across the whole codebase.
Access
Three ways to use it.
Your agent connects to a running server and drives the operations from there. We supply skills with worked examples, so the agent can write plans, chain them into flows, run them and keep them in order. Everything it builds appears in the workbench, where you inspect it.
Your engineers work hands-on: dependency views, dead-code triage, with the evidence behind each finding.
Your team writes and runs plans directly, against the same model.
In Python
What a plan looks like.
A first week with a legacy application, told as seven short scripts. The first three only read. The next two change the code. Each is checked or previewed before it is made. The last two look at interfaces and at how the application is divided into modules.
com.harbor.invoice, pricing, customer and ledger. Each script runs against a model of the whole program that the engine resolved beforehand, so a question about one method is answered from every file at once.What is in the codebase?
Every script starts from the same model of the program: its packages, types and methods, and every reference between them, already resolved.
query.packages lists the packages and how many types each holds. query.types lists the types that match a pattern. In a pattern, ** means anything below this point.
for package in query.packages(packageRestrictions=["com.harbor.**"]): print(package["primaryTypeCount"], package["packageName"]) for name in query.types(typeGlobs=["com.harbor.invoice.**"]): print(name)
Where are the longest methods?
methodMemberCounts measures every method: lines of code, complexity, parameters, the fields it touches. It can rank them on any of those. Here it returns the five longest in Harbor.
The pattern *(**) matches any method with any parameters. At the top of Harbor’s list is InvoiceService.calculateTotal, which the next steps work on.
longest = query.methodMemberCounts(methodGlobs=["com.harbor.**.*(**)"], rankMethodsBy=["linesOfCode"], limit=5) for method in longest: print(method["linesOfCode"], method["qualifiedMethod"])
What depends on it?
Before changing a method it helps to know who would feel it. query.callers returns every method that calls calculateTotal, anywhere in the codebase.
query.dependentTypes asks the same question one level up: every type that refers to InvoiceService in any way.
TOTAL = "com.harbor.invoice.InvoiceService.calculateTotal(**)" for caller in query.callers(methodGlobs=[TOTAL]): print(caller) users = query.dependentTypes(typeGlobs=["com.harbor.invoice.InvoiceService"]) print(len(users), "types refer to InvoiceService")
Check the rename, then make it.
The method adds tax, which its name does not say. With validationOnly set, rename.method runs every check the rename needs and writes nothing. A rename that cannot be made correctly raises a refusal instead of returning.
So the second call is only reached when the check passes. It renames the method and every call to it. The change is committed to the current branch, with a message written from what it did. except RefactorError catches a refusal and prints its reason. In that case nothing was written.
NEW_NAME = "totalIncludingTax" try: rename.method(targetMethods=[TOTAL], newMethod=NEW_NAME, validationOnly=True) renamed = rename.method(targetMethods=[TOTAL], newMethod=NEW_NAME) print(renamed["numberOfEdits"], "edits in", len(renamed["typesAffected"]), "types") print(renamed["suggestedCommitMessage"]) except RefactorError as refusal: print("refused:", refusal.msg, refusal.counts)
Break the long method into smaller ones.
splitMethodSuggestion proposes ways to cut a method into parts, each of which becomes a method of its own. Several scoring functions each make their own proposals. The script takes the highest-scoring proposal overall and hands it to splitMethod.
Some methods it will not split. If the same local variable name is declared in two different scopes, it refuses and names the variable to rename first.
METHOD = "com.harbor.invoice.InvoiceService.totalIncludingTax(com.harbor.invoice.Invoice)" proposals = extract.splitMethodSuggestion(qualifiedMethod=METHOD, targetLines=30) best = max((p for scorer in proposals for p in proposals[scorer]), key=lambda p: p["splitScore"]) print(len(best["newMethods"]), "new methods") extract.splitMethod(splitMethodSuggestion=best)
Put a narrower interface in front of the service.
Most classes that call InvoiceService use only a few of its methods. interfaceCandidates finds those few: the methods the callers actually call, and how many callers an interface of just those would serve.
extractInterfaceSuggestion then prices it before anything changes. It reports the interface it would create and what that would do to any dependency cycle the service is part of. Nothing is written until extractInterface is called.
ranked = graph.interfaceCandidates(targetTypes=["com.harbor.invoice.InvoiceService"], callerTypes=["com.harbor.**"]) best = ranked["rows"][0] proposal = extract.extractInterfaceSuggestion( fullyQualifiedName=best["targetType"], methodNames=best["methods"], destinationPackage="com.harbor.invoice.api", interfaceName="Invoicing") print(best["callersServed"], "callers could use", proposal["proposedInterface"]) print(proposal["cycleOutcome"])
Could pricing become a module of its own?
When two packages depend on each other, neither can be tested, reused or moved without the other. cyclesInPackageDependencyGraph returns each such cycle as its chain of dependencies. Harbor has one between invoice and pricing.
splitReadiness checks what would break if a set of types moved into a module of their own: imports left behind, packages split across both sides, build files that name them. Nothing is moved. For pricing the answer is no. Each message names something that holds it in place.
for cycle in graph.cyclesInPackageDependencyGraph(minCycleLength=2): print("cycle:", " -> ".join(edge["fromPackage"] for edge in cycle)) ready = graph.splitReadiness(typeGlobs=["com.harbor.pricing.**"], newModuleName="com.harbor.pricing") print("ready" if ready["splitIsReady"] else "not ready") for message in ready["messages"]: print(message["messageDescription"])
Hosting and requirements
On your hardware or in your cloud.
We love open source
Free up to half a million lines.
If your project is open source, the engine is free to use up to 500 KLOC. Get in touch and we will set you up.
