Quickstart
Install
python3 -m pip install tintelf
from tintelf import DeriveClient, DeriveFile
# Set DERIVE_API_KEY or pass token="<your token>"
derive = DeriveClient()
# Derive currently only supports CSV files.
obj = pd.read_csv("my_csv.csv")
file: DeriveFile = DeriveFile(content=obj, name="magic.csv")
graph_id, _ = derive.plan_graph(instructions=["Pass through the value of magic.csv"], data=[file])
results = derive.solve_graph(graph_id=graph_id, data=[file])
for x in results["objs"].keys():
print(derive.get_result(graph_id=graph_id, result_key=results["objs"][x][0]).content)
Results are retrieved as raw strings; it is up to the user to properly process results for later use (like using pd.read_csv).
In-Depth
Graphs
The core unit of Derive is the graph, specifying what operations are required, how they are configured, and what data they take as input to perform their calculations.
Graphs are created with plan_graph, and used with solve_graph. Graphs are composed of
models, which are blocks of operation that are connected together to form the full graph.
Presently the Derive team has included a set of models into Derive spanning
basic mathematics, statistics, correlation, and predictive analysis usecases.
Prompting
Prompts are lists of specific instructions, each instruction being referred to as a "step". Each step will be considered sequentially. Step specifications may reference the results of any previous step or the overall graph input as input to a given step.
# Example of referring to input and results
instructions = [
"Train a linear regression model on my_csv using val_1 and val_2 and targeting val_3",
"Use the trained regression model to predict based on val_4 and val_5 in my_other_csv"
]
Injecting Values
There may come a time when it is useful to inject a value into a given instruction, for
example when specifying an exponent where the value provided does not depend on graph input
or the results of previous steps. In these cases, Derive offers value injection, a clean
way to push static values into a graph. Values can be injected by writing your value between
square brackets in a prompt (e.g. [3]). Derive will automatically remember these values
and will refer to them by internal labels, like any other input or result in the graph.
Identical values will be treated as the same value, allowing you to reference the value
in multiple steps cleanly.
# Injecting values is simple!
instructions = [
"Multiply [12] by [53]",
"Multiply the result by [12]"
]
Instructions Derive cannot follow
Whenever Derive encounters an instruction that it cannot complete in one model, it will produce errors. Presently, the best solution for this is breaking the culprit instruction into two or more lower-level instructions. For example:
# Pythagorean theorem is not captured in Derive's model set;
# this will error.
instructions = [
"Solve the pythagorean theorem for all values of column_1, using column_1 as both A and B",
]
# Potential solution
instructions = [
"Square the values of column_1",
"Multiply the squared values by [2]",
"Apply square root to the summed values"
]
Addressing incorrect model/input selection
As of this writing, Derive retains some generative technologies within its core engine. Being in alpha, the Derive team is working to remove those elements and replace them with deterministic, analytical solutions. In the meantime, Derive may demonstrate hallmark behaviors of generative technologies, such as incorrectly identifying its own capabilities or failing to select the appropriate input or model.
Fortunately, there are steps which can be taken to alleviate this should it arise.
# Good
instructions = [
"Train a linear regression model on my_csv using val_1 and val_2 and targeting val_3",
"Use the trained regression model to predict based on val_4 and val_5 in my_other_csv"
]
# Bad
instructions = [
"Train with my_csv",
"Predict my_other_csv"
]